Forest vs Tree: The (N, K) Trade-off in Reproducible ML Evaluation
Deepak Pandita, Flip Korn, Chris Welty, Christopher M. Homan
摘要
Reproducibility is a cornerstone of scientific validation and of the authority it confers on its results. Reproducibility in machine learning evaluations leads to greater trust, confidence, and value. However, the ground truth responses used in machine learning often necessarily come from humans, among whom disagreement is prevalent, and surprisingly little research has studied the impact of effectively ignoring disagreement in these responses, as is typically the case. One reason for the lack of research is that budgets for collecting human-annotated evaluation data are limited, and obtaining more samples from multiple raters for each example greatly increases the per-item annotation costs. We investigate the trade-off between the number of items (N) and the number of responses per item (K) needed for reliable machine learning evaluation. We analyze a diverse collection of categorical datasets for which multiple annotations per item exist, and simulated distributions fit to these datasets, to determine the optimal (N, K) configuration, given a fixed budget (N x K), for collecting evaluation data and reliably comparing the performance of machine learning models. Our findings show, first, that accounting for human disagreement may come with N x K at no more than 1000 (and often much lower) for every dataset tested on at least one metric. Moreover, this minimal N x K almost always occurred for K > 10. Furthermore, the nature of the tradeoff between K and N, or if one even existed, depends on the evaluation metric, with metrics that are more sensitive to the full distribution of responses performing better at higher levels of K. Our methods can be used to help ML practitioners get more effective test data by finding the optimal metrics and number of items and annotations per item to collect to get the most reliability for their budget.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Toward a Perspectivist Turn in Ground Truthing for Predictive ComputingFederico Cabitza, Andrea Campagner, Valerio BasileAAAI 2023 · 被引用 236 次
- Reproducibility in Computational Linguistics: Is Source Code Enough?Mohammad Arvan, Luís Pina, Natalie PardeEMNLP 2022 · 被引用 12 次
- Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is OffensiveTharindu Cyril Weerasooriya, Sujan Dutta, Tharindu Ranasinghe, Marcos Zampieri 等EMNLP 2023 · 被引用 12 次
- D3CODE: Disentangling Disagreements in Data across Cultures on Offensiveness Detection and EvaluationAida Mostafazadeh Davani, Mark Diaz, Dylan K. Baker, Vinodkumar PrabhakaranEMNLP 2024 · 被引用 3 次
相关 Paper
- Discrepancy Ratio: Evaluating Model Performance When Even Experts Disagree on the TruthIgor Lovchinsky, Alon Daks, Israel Malkin, Pouya Samangouei 等ICLR 2020 · 被引用 11 次
- Towards Inferential Reproducibility of Machine Learning ResearchMichael Hagmann, Philipp Meier, Stefan RiezlerICLR 2023 · 被引用 1 次
- Validating LLM-as-a-Judge Systems under Rating IndeterminacyLuke Guerdan, Solon Barocas, Kenneth Holstein, Hanna M. Wallach 等NeurIPS 2025 · 被引用 31 次
- NUTMEG: Separating Signal From Noise in Annotator DisagreementJonathan Ivey, Susan Gauch, David JurgensEMNLP 2025
- Reliance and Automation for Human-AI Collaborative Data Labeling Conflict ResolutionMichelle Brachman, Zahra Ashktorab, Michael Desmond, Evelyn Duesterwald 等CSCW 2022 · 被引用 16 次
