The Challenge of Variable Effort Crowdsourcing and How Visible Gold Can Help
Danula Hettiachchi, Mike Schaekermann, Tristan McKinney, Matthew Lease
摘要
We consider a class of variable effort human annotation tasks in which the number of labels required per item can greatly vary (e.g., finding all faces in an image, named entities in a text, bird calls in an audio recording, etc.). In such tasks, some items require far more effort than others to annotate. Furthermore, the per-item annotation effort is not known until after each item is annotated since determining the number of labels required is an implicit part of the annotation task itself. On an image bounding-box task with crowdsourced annotators, we show that annotator accuracy and recall consistently drop as effort increases. We hypothesize reasons for this drop and investigate a set of approaches to counteract it. Firstly, we benchmark on this task a set of general best-practice methods for quality crowdsourcing. Notably, only one of these methods actually improves quality: the use of visible gold questions that provide periodic feedback to workers on their accuracy as they work. Given these promising results, we then investigate and evaluate variants of the visible gold approach, yielding further improvement. Final results show a 7% improvement in bounding-box accuracy over the baseline. We discuss the generality of the visible gold approach and promising directions for future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Prompting in the Dark: Assessing Human Performance in Prompt Engineering for Data Labeling When Gold Labels Are AbsentZeyu He, Saniya Naphade, Ting-Hao 'Kenneth' HuangCHI 2025 · 被引用 22 次
- LabelAId: Just-in-time AI Interventions for Improving Human Labeling Quality and Domain Knowledge in Crowdsourcing SystemsChu Li, Zhihan Zhang, Michael Saugstad, Esteban Safranchik 等CHI 2024 · 被引用 12 次
- Judgment Sieve: Reducing Uncertainty in Group Judgments through Interventions Targeting Ambiguity versus DisagreementQuan Ze Chen, Amy X. ZhangCSCW 2023 · 被引用 10 次
- Voices in a Crowd: Searching for clusters of unique perspectivesNikolas Vitsakis, Amit Parekh, Ioannis KonstasEMNLP 2024
它引用的顶会 Paper4
- Objects365: A Large-Scale, High-Quality Dataset for Object DetectionShuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng 等ICCV 2019 · 被引用 1,018 次
- Becoming the Super Turker: Increasing Wages via a Strategy from High Earning WorkersSaiph Savage, Chun-Wei Chiang, Susumu Saito, Carlos Toxtli 等WWW 2020 · 被引用 53 次
- Ambiguity-aware AI Assistants for Medical Data AnalysisMike Schaekermann, Graeme Beaton, Elaheh Sanoubari, Andrew Lim 等CHI 2020 · 被引用 45 次
- CrowdCog: A Cognitive Skill based System for Heterogeneous Task Assignment and Recommendation in CrowdsourcingDanula Hettiachchi, Niels van Berkel, Vassilis Kostakos, Jorge GonçalvesCSCW 2020 · 被引用 38 次
相关 Paper
- Efficient Online Crowdsourcing with Complex AnnotationsReshef Meir, Viet-An Nguyen, Xu Chen, Jagdish Ramakrishnan 等AAAI 2024 · 被引用 1 次
- Detecting and Preventing Confused Labels in Crowdsourced DataEvgeny Krivosheev, Siarhei Bykau, Fabio Casati, Sunil PrabhakarVLDB 2020 · 被引用 12 次
- Towards Good Practices for Efficiently Annotating Large-Scale Image Classification DatasetsYuan-Hong Liao, Amlan Kar, Sanja FidlerCVPR 2021
- Dynamic Compensation Can Enhance User Engagement by Triggering Sensitivity to Financial Losses in Crowd-sourced StudiesCatalina Gomez, Mung Yao Jia, Sue Min Cho, Chien-Ming Huang 等CHI 2026 · 被引用 1 次
- Crowdsourcing Learning as Domain Adaptation: A Case Study on Named Entity RecognitionXin Zhang, Guangwei Xu, Yueheng Sun, Meishan Zhang 等ACL 2021
