Goldilocks: Consistent Crowdsourced Scalar Annotations with Relative Uncertainty
Quanze Chen, Daniel S. Weld, Amy X. Zhang
Abstract
Human ratings have become a crucial resource for training and evaluating machine learning systems. However, traditional elicitation methods for absolute and comparative rating suffer from issues with consistency and often do not distinguish between uncertainty due to disagreement between annotators and ambiguity inherent to the item being rated. In this work, we present Goldilocks, a novel crowd rating elicitation technique for collecting calibrated scalar annotations that also distinguishes inherent ambiguity from inter-annotator disagreement. We introduce two main ideas: grounding absolute rating scales with examples and using a two-step bounding process to establish a range for an item's placement. We test our designs in three domains: judging toxicity of online comments, estimating satiety of food depicted in images, and estimating age based on portraits. We show that (1) Goldilocks can improve consistency in domains where interpretation of the scale is not universal, and that (2) representing items with ranges lets us simultaneously capture different sources of uncertainty leading to better estimates of pairwise relationship distributions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a4086ee-eb5f-41ee-aadc-ef36f77dc8a7Cited by top-tier papers5
- Judgment Sieve: Reducing Uncertainty in Group Judgments through Interventions Targeting Ambiguity versus DisagreementQuan Ze Chen, Amy X. ZhangCSCW 2023 · 10 citations
- Are Human Explanations Always Helpful? Towards Objective Evaluation of Human Natural Language ExplanationsBingsheng Yao, Prithviraj Sen, Lucian Popa, James A. Hendler et al.ACL 2023 · 4 citations
- Wisdom of the Crowd, Without the Crowd: A Socratic LLM for Asynchronous Deliberation on Perspectivist DataMalik Khadar, Daniel Runningen, Julia Tang, Stevie Chancellor et al.CSCW 2025 · 2 citations
- Promptimizer: User-Led Prompt Optimization for Personal Content ClassificationLeijie Wang, Kathryn Yurechko, Amy X. ZhangCHI 2026 · 1 citation
- Measuring scalar constructs in social science with LLMsHauke Licht, Rupak Sarkar, Patrick Y. Wu, Pranav Goel et al.EMNLP 2025
Builds on4
- The Disagreement Deconvolution: Bringing Machine Learning Performance Metrics In Line With RealityMitchell L. Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto et al.CHI 2021 · 100 citations
- Beyond User Self-Reported Likert Scale Ratings: A Comparison Model for Automatic Dialog EvaluationWeixin Liang, James Zou, Zhou YuACL 2020 · 25 citations
- Rank Aggregation via Heterogeneous Thurstone Preference ModelsTao Jin, Pan Xu, Quanquan Gu, Farzad FarnoudAAAI 2020 · 19 citations
- Crowdsourced Detection of Emotionally Manipulative LanguageJordan S. Huffaker, Jonathan K. Kummerfeld, Walter S. Lasecki, Mark S. AckermanCHI 2020 · 9 citations
Related papers
- Validating LLM-as-a-Judge Systems under Rating IndeterminacyLuke Guerdan, Solon Barocas, Kenneth Holstein, Hanna M. Wallach et al.NeurIPS 2025 · 31 citations
- AtC: Aggregate-then-Calibrate for Human-centered AssessmentZejun Xie, Xintong Li, Guang Wang, Desheng ZhangICLR 2026
- The Challenge of Variable Effort Crowdsourcing and How Visible Gold Can HelpDanula Hettiachchi, Mike Schaekermann, Tristan McKinney, Matthew LeaseCSCW 2021 · 21 citations
- Eliciting Confidence for Improving Crowdsourced Audio AnnotationsAna Elisa Méndez Méndez, Mark Cartwright, Juan Pablo Bello, Oded NovCSCW 2022 · 10 citations
- Soft-Label Integration for Robust Toxicity ClassificationZelei Cheng, Xian Wu, Jiahao Yu, Shuo Han et al.NeurIPS 2024 · 7 citations
