On Efficient and Statistical Quality Estimation for Data Annotation
Jan-Christoph Klie, Juan Haladjian, Marc Kirchner, Rahul Nair
摘要
Annotated datasets are an essential ingredient to train, evaluate, compare and productionalize supervised machine learning models. It is therefore imperative that annotations are of high quality. For their creation, good quality management and thereby reliable quality estimates are needed. Then, if quality is insufficient during the annotation process, rectifying measures can be taken to improve it. Quality estimation is often performed by having experts manually label instances as correct or incorrect. But checking all annotated instances tends to be expensive. Therefore, in practice, usually only subsets are inspected; sizes are chosen mostly without justification or regard to statistical power and more often than not, are relatively small. Basing estimates on small sample sizes, however, can lead to imprecise values for the error rate. Using unnecessarily large sample sizes costs money that could be better spent, for instance on more annotations. Therefore, we first describe in detail how to use confidence intervals for finding the minimal sample size needed to estimate the annotation error rate. Then, we propose applying acceptance sampling as an alternative to error rate estimation We show that acceptance sampling can reduce the required sample sizes up to 50% while providing the same statistical guarantees.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong 等CHI 2021 · 被引用 725 次
- When does dough become a bagel? Analyzing the remaining mistakes on ImageNetVijay Vasudevan, Benjamin Caine, Raphael Gontijo Lopes, Sara Fridovich-Keil 等NeurIPS 2022 · 被引用 79 次
相关 Paper
- Learning a Cost-Effective Annotation Policy for Question AnsweringBernhard Kratzwald, Stefan Feuerriegel, Huan SunEMNLP 2020 · 被引用 9 次
- Promises and Pitfalls of Threshold-based Auto-labelingHarit Vishwakarma, Heguang Lin, Frederic Sala, Ramya Korlakai VinayakNeurIPS 2023 · 被引用 16 次
- Low-Shot Validation: Active Importance Sampling for Estimating Classifier Performance on Rare CategoriesFait Poms, Vishnu Sarukkai, Ravi Teja Mullapudi, Nimit Sharad Sohoni 等ICCV 2021 · 被引用 10 次
- A Needle in a Haystack: An Analysis of High-Agreement Workers on MTurk for SummarizationLining Zhang, Simon Mille, Yufang Hou, Daniel Deutsch 等ACL 2023 · 被引用 6 次
- Forest vs Tree: The (N, K) Trade-off in Reproducible ML EvaluationDeepak Pandita, Flip Korn, Chris Welty, Christopher M. HomanAAAI 2026 · 被引用 2 次
