Crowdsourcing enables robust cell annotation for breast cancer pathology
Pascal Klöckner (University of Bern, Netherlands Cancer Institute, University Hospital of Psychiatry, ChangeGamers)
Melis Erdal Cesur (Netherlands Cancer Institute)
Bart de Rooij (Netherlands Cancer Institute)
Zainab Ameziane (Netherlands Cancer Institute)
Idris Iritas (Netherlands Cancer Institute)
Rolf Harkes (Netherlands Cancer Institute)
Rafael Bidarra (TU Delft - Electrical Engineering, Mathematics and Computer Science)
Sara Pires Oliveira (Netherlands Cancer Institute)
Hugo Mark Horlings (ChangeGamers, Netherlands Cancer Institute)
More Authors
More Info
expand_more
Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.
Abstract
The tumor microenvironment (TME) contains important morphological and molecular cues that help determine prognosis and therapeutic responses. Deep learning models can assist pathologists in assessing such biomarkers. However, generating high-quality annotations for computational pathology remains a significant bottleneck due to the level of expertise and considerable time required for fine-grained labeling. In this study, we evaluate the feasibility of using non-expert crowdsourcing for cell annotation in hematoxylin and eosin (H&E) samples of breast cancer tissue. Unlike prior work, performance assessment was based on high-fidelity ground-truth labels derived from multiplexed immunofluorescence using CODEX data, allowing for accurate benchmarking of both experts and non-experts. We collected cell annotations from experts, semi-experts, and non-experts through Tilly, a gamified annotation application designed to train and engage users in identifying major cell types within the TME. Overall, our results show that non-expert crowdsourcing is a scalable and effective strategy for generating training data for the classification of major cell types in H&E images: tumor cells, lymphocytes, and fibroblasts. Moreover, combining large, crowdsourced datasets with smaller, high-quality subsets annotated using spatial proteomics may offer a practical annotation approach to develop more robust models while minimizing biases.