The Disagreement Deconvolution: Bringing Machine Learning Performance Metrics In Line With Reality
Mitchell L. Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto, Michael S. Bernstein
摘要
Machine learning classifiers for human-facing tasks such as comment toxicity and misinformation often score highly on metrics such as ROC AUC but are received poorly in practice. Why this gap? Today, metrics such as ROC AUC, precision, and recall are used to measure technical performance; however, human-computer interaction observes that evaluation of human-facing systems should account for people’s reactions to the system. In this paper, we introduce a transformation that more closely aligns machine learning classification metrics with the values and methods of user-facing performance measures. The disagreement deconvolution takes in any multi-annotator (e.g., crowdsourced) dataset, disentangles stable opinions from noise by estimating intra-annotator consistency, and compares each test set prediction to the individual stable opinions from each annotator. Applying the disagreement deconvolution to existing social computing datasets, we find that current metrics dramatically overstate the performance of many human-facing machine learning tasks: for example, performance on a comment toxicity task is corrected from .95 to .73 ROC AUC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper41
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee 等ICML 2023 · 被引用 764 次
- Self-Supervised Contrastive Pre-Training For Time Series via Time-Frequency ConsistencyXiang Zhang, Ziyuan Zhao, Theodoros Tsiligkaridis, Marinka ZitnikNeurIPS 2022 · 被引用 558 次
- Explanations Can Reduce Overreliance on AI Systems During Decision-MakingHelena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg 等CSCW 2023 · 被引用 362 次
- FLASK: Fine-grained Language Model Evaluation based on Alignment Skill SetsSeonghyeon Ye, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang 等ICLR 2024 · 被引用 176 次
- Trauma-Informed Computing: Towards Safer Technology Experiences for AllJanet X. Chen, Allison McDonald, Yixin Zou, Emily Tseng 等CHI 2022 · 被引用 163 次
它引用的顶会 Paper3
- Human Uncertainty Makes Classification More RobustJoshua C. Peterson, Ruairidh M. Battleday, Thomas L. Griffiths, Olga RussakovskyICCV 2019 · 被引用 362 次
- Designing Ground Truth and the Social Life of LabelsMichael J. Muller, Christine T. Wolf, Josh Andres, Michael Desmond 等CHI 2021 · 被引用 80 次
- Learning from Crowds by Modeling Common ConfusionsZhendong Chu, Jing Ma, Hongning WangAAAI 2021 · 被引用 60 次
相关 Paper
- Discrepancy Ratio: Evaluating Model Performance When Even Experts Disagree on the TruthIgor Lovchinsky, Alon Daks, Israel Malkin, Pouya Samangouei 等ICLR 2020 · 被引用 11 次
- STABLEVAL: Disagreement-Aware and Stable Evaluation of AI SystemsSailendra Akash Bonagiri, Gerard Anderias, Saee Patil, Angelina Lai 等ICML 2026 · 被引用 2 次
- Forest vs Tree: The (N, K) Trade-off in Reproducible ML EvaluationDeepak Pandita, Flip Korn, Chris Welty, Christopher M. HomanAAAI 2026 · 被引用 2 次
- When the Majority is Wrong: Modeling Annotator Disagreement for Subjective TasksEve Fleisig, Rediet Abebe, Dan KleinEMNLP 2023 · 被引用 11 次
- Evaluating the Interpretability of Generative Models by Interactive ReconstructionAndrew Slavin Ross, Nina Chen, Elisa Zhao Hang, Elena L. Glassman 等CHI 2021 · 被引用 40 次
