Interface Design for Crowdsourcing Hierarchical Multi-Label Text Annotations
Rickard Stureborg, Bhuwan Dhingra, Jun Yang
Abstract
Human data labeling is an important and expensive task at the heart of supervised learning systems. Hierarchies help humans understand and organize concepts. We ask whether and how concept hierarchies can inform the design of annotation interfaces to improve labeling quality and efficiency. We study this question through annotation of vaccine misinformation, where the labeling task is difficult and highly subjective. We investigate 6 user interface designs for crowdsourcing hierarchical labels by collecting over 18,000 individual annotations. Under a fixed budget, integrating hierarchies into the design improves crowdsource workers' F1 scores. We attribute this to (1) Grouping similar concepts, improving F1 scores by +0.16 over random groupings, (2) Strong relative performance on high-difficulty examples (relative F1 score difference of +0.40), and (3) Filtering out obvious negatives, increasing precision by +0.07. Ultimately, labeling schemes integrating the hierarchy outperform those that do not -achieving mean F1 of 0.70.
• Human-centered computing → HCI design and evaluation methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7975b213-ba27-4cc7-99e4-a95e1a17aa3cCited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- Is this AI trained on Credible Data? The Effects of Labeling Quality and Performance Bias on User TrustCheng Chen, S. Shyam SundarCHI 2023 · 36 citations
- Hierarchical Crowdsourcing for Data Labeling with Heterogeneous CrowdHaodi Zhang, Wenxi Huang, Zhenhan Su, Junyang Chen et al.ICDE 2023 · 4 citations
- Understanding Human-side Impact of Sampling Image Batches in Subjective Attribute LabelingChaeyeon Chung, Jungsoo Lee, Kyungmin Park, Junsoo Lee et al.CSCW 2021 · 4 citations
- Investigating Differences in Crowdsourced News Credibility Assessment: Raters, Tasks, and Expert CriteriaMd Momen Bhuiyan, Amy X. Zhang, Connie Moon Sehat, Tanushree MitraCSCW 2020 · 70 citations
- Crowd Teaching with Imperfect LabelsYao Zhou, Arun Reddy Nelakurthi, Ross Maciejewski, Wei Fan et al.WWW 2020 · 12 citations
