The Challenge of Variable Effort Crowdsourcing and How Visible Gold Can Help
Danula Hettiachchi, Mike Schaekermann, Tristan McKinney, Matthew Lease
Abstract
We consider a class of variable effort human annotation tasks in which the number of labels required per item can greatly vary (e.g., finding all faces in an image, named entities in a text, bird calls in an audio recording, etc.). In such tasks, some items require far more effort than others to annotate. Furthermore, the per-item annotation effort is not known until after each item is annotated since determining the number of labels required is an implicit part of the annotation task itself. On an image bounding-box task with crowdsourced annotators, we show that annotator accuracy and recall consistently drop as effort increases. We hypothesize reasons for this drop and investigate a set of approaches to counteract it. Firstly, we benchmark on this task a set of general best-practice methods for quality crowdsourcing. Notably, only one of these methods actually improves quality: the use of visible gold questions that provide periodic feedback to workers on their accuracy as they work. Given these promising results, we then investigate and evaluate variants of the visible gold approach, yielding further improvement. Final results show a 7% improvement in bounding-box accuracy over the baseline. We discuss the generality of the visible gold approach and promising directions for future research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e92523ee-7e74-4a73-bbbf-d6e7e293eceaCited by top-tier papers4
- Prompting in the Dark: Assessing Human Performance in Prompt Engineering for Data Labeling When Gold Labels Are AbsentZeyu He, Saniya Naphade, Ting-Hao 'Kenneth' HuangCHI 2025 · 22 citations
- LabelAId: Just-in-time AI Interventions for Improving Human Labeling Quality and Domain Knowledge in Crowdsourcing SystemsChu Li, Zhihan Zhang, Michael Saugstad, Esteban Safranchik et al.CHI 2024 · 12 citations
- Judgment Sieve: Reducing Uncertainty in Group Judgments through Interventions Targeting Ambiguity versus DisagreementQuan Ze Chen, Amy X. ZhangCSCW 2023 · 10 citations
- Voices in a Crowd: Searching for clusters of unique perspectivesNikolas Vitsakis, Amit Parekh, Ioannis KonstasEMNLP 2024
Builds on4
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng et al.ICCV 2019 · 1,018 citations
- Becoming the Super Turker: Increasing Wages via a Strategy from High Earning WorkersSaiph Savage, Chun-Wei Chiang, Susumu Saito, Carlos Toxtli et al.WWW 2020 · 53 citations
- Ambiguity-aware AI Assistants for Medical Data AnalysisMike Schaekermann, Graeme Beaton, Elaheh Sanoubari, Andrew Lim et al.CHI 2020 · 45 citations
- CrowdCog: A Cognitive Skill based System for Heterogeneous Task Assignment and Recommendation in CrowdsourcingDanula Hettiachchi, Niels van Berkel, Vassilis Kostakos, Jorge GonçalvesCSCW 2020 · 38 citations
Related papers
- Efficient Online Crowdsourcing with Complex AnnotationsReshef Meir, Viet-An Nguyen, Xu Chen, Jagdish Ramakrishnan et al.AAAI 2024 · 1 citation
- Detecting and Preventing Confused Labels in Crowdsourced DataEvgeny Krivosheev, Siarhei Bykau, Fabio Casati, Sunil PrabhakarVLDB 2020 · 12 citations
- Towards Good Practices for Efficiently Annotating Large-Scale Image Classification DatasetsYuan-Hong Liao, Amlan Kar, Sanja FidlerCVPR 2021
- Dynamic Compensation Can Enhance User Engagement by Triggering Sensitivity to Financial Losses in Crowd-sourced StudiesCatalina Gomez, Mung Yao Jia, Sue Min Cho, Chien-Ming Huang et al.CHI 2026 · 1 citation
- Crowdsourcing Learning as Domain Adaptation: A Case Study on Named Entity RecognitionXin Zhang, Guangwei Xu, Yueheng Sun, Meishan Zhang et al.ACL 2021
