Goldilocks: Consistent Crowdsourced Scalar Annotations with Relative Uncertainty
Quanze Chen, Daniel S. Weld, Amy X. Zhang
摘要
Human ratings have become a crucial resource for training and evaluating machine learning systems. However, traditional elicitation methods for absolute and comparative rating suffer from issues with consistency and often do not distinguish between uncertainty due to disagreement between annotators and ambiguity inherent to the item being rated. In this work, we present Goldilocks, a novel crowd rating elicitation technique for collecting calibrated scalar annotations that also distinguishes inherent ambiguity from inter-annotator disagreement. We introduce two main ideas: grounding absolute rating scales with examples and using a two-step bounding process to establish a range for an item's placement. We test our designs in three domains: judging toxicity of online comments, estimating satiety of food depicted in images, and estimating age based on portraits. We show that (1) Goldilocks can improve consistency in domains where interpretation of the scale is not universal, and that (2) representing items with ranges lets us simultaneously capture different sources of uncertainty leading to better estimates of pairwise relationship distributions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Judgment Sieve: Reducing Uncertainty in Group Judgments through Interventions Targeting Ambiguity versus DisagreementQuan Ze Chen, Amy X. ZhangCSCW 2023 · 被引用 10 次
- Are Human Explanations Always Helpful? Towards Objective Evaluation of Human Natural Language ExplanationsBingsheng Yao, Prithviraj Sen, Lucian Popa, James A. Hendler 等ACL 2023 · 被引用 4 次
- Wisdom of the Crowd, Without the Crowd: A Socratic LLM for Asynchronous Deliberation on Perspectivist DataMalik Khadar, Daniel Runningen, Julia Tang, Stevie Chancellor 等CSCW 2025 · 被引用 2 次
- Promptimizer: User-Led Prompt Optimization for Personal Content ClassificationLeijie Wang, Kathryn Yurechko, Amy X. ZhangCHI 2026 · 被引用 1 次
- Measuring scalar constructs in social science with LLMsHauke Licht, Rupak Sarkar, Patrick Y. Wu, Pranav Goel 等EMNLP 2025
它引用的顶会 Paper4
- The Disagreement Deconvolution: Bringing Machine Learning Performance Metrics In Line With RealityMitchell L. Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto 等CHI 2021 · 被引用 100 次
- Beyond User Self-Reported Likert Scale Ratings: A Comparison Model for Automatic Dialog EvaluationWeixin Liang, James Zou, Zhou YuACL 2020 · 被引用 25 次
- Rank Aggregation via Heterogeneous Thurstone Preference ModelsTao Jin, Pan Xu, Quanquan Gu, Farzad FarnoudAAAI 2020 · 被引用 19 次
- Crowdsourced Detection of Emotionally Manipulative LanguageJordan S. Huffaker, Jonathan K. Kummerfeld, Walter S. Lasecki, Mark S. AckermanCHI 2020 · 被引用 9 次
相关 Paper
- Validating LLM-as-a-Judge Systems under Rating IndeterminacyLuke Guerdan, Solon Barocas, Kenneth Holstein, Hanna M. Wallach 等NeurIPS 2025 · 被引用 31 次
- AtC: Aggregate-then-Calibrate for Human-centered AssessmentZejun Xie, Xintong Li, Guang Wang, Desheng ZhangICLR 2026
- The Challenge of Variable Effort Crowdsourcing and How Visible Gold Can HelpDanula Hettiachchi, Mike Schaekermann, Tristan McKinney, Matthew LeaseCSCW 2021 · 被引用 21 次
- Eliciting Confidence for Improving Crowdsourced Audio AnnotationsAna Elisa Méndez Méndez, Mark Cartwright, Juan Pablo Bello, Oded NovCSCW 2022 · 被引用 10 次
- Soft-Label Integration for Robust Toxicity ClassificationZelei Cheng, Xian Wu, Jiahao Yu, Shuo Han 等NeurIPS 2024 · 被引用 7 次
