Eliciting Confidence for Improving Crowdsourced Audio Annotations
Ana Elisa Méndez Méndez, Mark Cartwright, Juan Pablo Bello, Oded Nov
摘要
In this work we explore confidence elicitation methods for crowdsourcing "soft" labels, e.g., probability estimates, to reduce the annotation costs for domains with ambiguous data. Machine learning research has shown that such "soft" labels are more informative and can reduce the data requirements when training supervised machine learning models. By reducing the number of required labels, we can reduce the costs of slow annotation processes such as audio annotation. In our experiments we evaluated three confidence elicitation methods: 1) "No Confidence" elicitation, 2) "Simple Confidence" elicitation, and 3) "Betting" mechanism for confidence elicitation, at both individual (i.e., per participant) and aggregate (i.e., crowd) levels. In addition, we evaluated the interaction between confidence elicitation methods, annotation types (binary, probability, and z-score derived probability), and "soft" versus "hard" (i.e., binarized) aggregate labels. Our results show that both confidence elicitation mechanisms result in higher annotation quality than the "No Confidence" mechanism for binary annotations at both participant and recording levels. In addition, when aggregating labels at the recording level, results indicate that we can achieve comparable results to those with 10-participant aggregate annotations using fewer annotators if we aggregate "soft" labels instead of "hard" labels. These results suggest that for binary audio annotation using a confidence elicitation mechanism and aggregating continuous labels we can obtain higher annotation quality, more informative labels, with quality differences more pronounced with fewer participants. Finally, we propose a way of integrating these confidence elicitation methods into a two-stage, multi-label annotation pipeline.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Recovering Top-Two Answers and Confusion Probability in Multi-Choice CrowdsourcingHyeonsu Jeong, Hye Won ChungICML 2023 · 被引用 2 次
- A Probabilistic Graphical Model for Analyzing the Subjective Visual Quality Assessment Data from CrowdsourcingJing Li, Suiyi Ling, Junle Wang, Patrick Le CalletACM MM 2020 · 被引用 23 次
- MAS: Model-Agnostic Active Annotation Strategy for CrowdsourcingWenjun Zhang, Liangxiao Jiang, Chaoqun Li, Shanshan SiICML 2026
- Crowd Teaching with Imperfect LabelsYao Zhou, Arun Reddy Nelakurthi, Ross Maciejewski, Wei Fan 等WWW 2020 · 被引用 12 次
- Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated InterpretabilityTuomas Oikarinen, Ge Yan, Akshay Kulkarni, Tsui-Wei WengCVPR 2026 · 被引用 1 次
