Spot Check Equivalence: An Interpretable Metric for Information Elicitation Mechanisms
Shengwei Xu, Yichi Zhang, Paul Resnick, Grant Schoenebeck
Abstract
Because high-quality data is like oxygen for AI systems, effectively eliciting information from crowdsourcing workers has become a first-order problem for developing high-performance machine learning algorithms. Two prevalent paradigms, spot-checking and peer prediction, enable the design of mechanisms to evaluate and incentivize high-quality data from human labelers. So far, at least three metrics have been proposed to compare the performances of these techniques 2022high,gao2016incentivizing,burrell2021measurement. However, different metrics lead to divergent and even contradictory results in various contexts. In this paper, we harmonize these divergent stories, showing that two of these metrics are actually the same within certain contexts and explain the divergence of the third. Moreover, we unify these different contexts by introducingSpot Check Equivalence, which offers an interpretable metric for the effectiveness of a peer prediction mechanism. Finally, we present two approaches to compute spot check equivalence in various contexts, where simulation results verify the effectiveness of our proposed metric.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Stochastically Dominant Peer PredictionYichi Zhang, Shengwei Xu, Grant Schoenebeck, David M. PennockNeurIPS 2025 · 2 citations
- Benchmarking LLMs' Judgments with No Gold StandardShengwei Xu, Yuxuan Lu, Grant Schoenebeck, Yuqing KongICLR 2025
Builds on5
- Dominantly Truthful Multi-task Peer Prediction with a Constant Number of TasksYuqing KongSODA 2020 · 33 citations
- Information Elicitation from Rowdy CrowdsGrant Schoenebeck, Fang-Yi Yu, Yichi ZhangWWW 2021 · 18 citations
- High-Effort Crowds: Limited Liability via TournamentsYichi Zhang, Grant SchoenebeckWWW 2023 · 10 citations
- Eliciting Thinking Hierarchy without a PriorYuqing Kong, Yunqi Li, Yubo Zhang, Zhihuan Huang et al.NeurIPS 2022 · 9 citations
- Multitask Peer Prediction With Task-dependent StrategiesYichi Zhang, Grant SchoenebeckWWW 2023 · 7 citations
Related papers
- Evaluating LLM-contaminated Crowdsourcing Data Without Ground TruthYichi Zhang, Jinlong Pang, Zhaowei Zhu, Yang LiuNeurIPS 2025 · 3 citations
- HybridEval: A Human-AI Collaborative Approach for Evaluating Design Ideas at ScaleSepideh Mesbah, Ines Arous, Jie Yang, Alessandro BozzonWWW 2023 · 5 citations
- What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks?Nikita Nangia, Saku Sugawara, Harsh Trivedi, Alex Warstadt et al.ACL 2021
- CrowdCO-OP: Sharing Risks and Rewards in CrowdsourcingShaoyang Fan, Ujwal Gadiraju, Alessandro Checco, Gianluca DemartiniCSCW 2020 · 34 citations
- Peer Neighborhood Mechanisms: A Framework for Mechanism GeneralizationAdam Richardson, Boi FaltingsAAAI 2024 · 1 citation
