The Disagreement Deconvolution: Bringing Machine Learning Performance Metrics In Line With Reality
Mitchell L. Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto, Michael S. Bernstein
Abstract
Machine learning classifiers for human-facing tasks such as comment toxicity and misinformation often score highly on metrics such as ROC AUC but are received poorly in practice. Why this gap? Today, metrics such as ROC AUC, precision, and recall are used to measure technical performance; however, human-computer interaction observes that evaluation of human-facing systems should account for people’s reactions to the system. In this paper, we introduce a transformation that more closely aligns machine learning classification metrics with the values and methods of user-facing performance measures. The disagreement deconvolution takes in any multi-annotator (e.g., crowdsourced) dataset, disentangles stable opinions from noise by estimating intra-annotator consistency, and compares each test set prediction to the individual stable opinions from each annotator. Applying the disagreement deconvolution to existing social computing datasets, we find that current metrics dramatically overstate the performance of many human-facing machine learning tasks: for example, performance on a comment toxicity task is corrected from .95 to .73 ROC AUC.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b5f4a96-2e23-4dc8-998c-a0db53c1b46dCited by top-tier papers41
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee et al.ICML 2023 · 764 citations
- Self-Supervised Contrastive Pre-Training For Time Series via Time-Frequency ConsistencyXiang Zhang, Ziyuan Zhao, Theodoros Tsiligkaridis, Marinka ZitnikNeurIPS 2022 · 558 citations
- Explanations Can Reduce Overreliance on AI Systems During Decision-MakingHelena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg et al.CSCW 2023 · 362 citations
- FLASK: Fine-grained Language Model Evaluation based on Alignment Skill SetsSeonghyeon Ye, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang et al.ICLR 2024 · 176 citations
- Trauma-Informed Computing: Towards Safer Technology Experiences for AllJanet X. Chen, Allison McDonald, Yixin Zou, Emily Tseng et al.CHI 2022 · 163 citations
Builds on3
- Human Uncertainty Makes Classification More RobustJoshua C. Peterson, Ruairidh M. Battleday, Thomas L. Griffiths, Olga RussakovskyICCV 2019 · 362 citations
- Designing Ground Truth and the Social Life of LabelsMichael J. Muller, Christine T. Wolf, Josh Andres, Michael Desmond et al.CHI 2021 · 80 citations
- Learning from Crowds by Modeling Common ConfusionsZhendong Chu, Jing Ma, Hongning WangAAAI 2021 · 60 citations
Related papers
- Discrepancy Ratio: Evaluating Model Performance When Even Experts Disagree on the TruthIgor Lovchinsky, Alon Daks, Israel Malkin, Pouya Samangouei et al.ICLR 2020 · 11 citations
- STABLEVAL: Disagreement-Aware and Stable Evaluation of AI SystemsSailendra Akash Bonagiri, Gerard Anderias, Saee Patil, Angelina Lai et al.ICML 2026 · 2 citations
- Forest vs Tree: The (N, K) Trade-off in Reproducible ML EvaluationDeepak Pandita, Flip Korn, Chris Welty, Christopher M. HomanAAAI 2026 · 2 citations
- When the Majority is Wrong: Modeling Annotator Disagreement for Subjective TasksEve Fleisig, Rediet Abebe, Dan KleinEMNLP 2023 · 11 citations
- Evaluating the Interpretability of Generative Models by Interactive ReconstructionAndrew Slavin Ross, Nina Chen, Elisa Zhao Hang, Elena L. Glassman et al.CHI 2021 · 40 citations
