Discrepancy Ratio: Evaluating Model Performance When Even Experts Disagree on the Truth
Igor Lovchinsky, Alon Daks, Israel Malkin, Pouya Samangouei, Ardavan Saeedi, Yang Liu, Swami Sankaranarayanan, Tomer Gafner, Ben Sternlieb, Patrick Maher, Nathan Silberman
Abstract
In most machine learning tasks unambiguous ground truth labels can easily be acquired. However, this luxury is often not afforded to many high-stakes, real-world scenarios such as medical image interpretation, where even expert human annotators typically exhibit very high levels of disagreement with one another. While prior works have focused on overcoming noisy labels during training, the question of how to evaluate models when annotators disagree about ground truth has remained largely unexplored. To address this, we propose the discrepancy ratio: a novel, task-independent and principled framework for validating machine learning models in the presence of high label noise. Conceptually, our approach evaluates a model by comparing its predictions to those of human annotators, taking into account the degree to which annotators disagree with one another. While our approach is entirely general, we show that in the special case of binary classification, our proposed metric can be evaluated in terms of simple, closed-form expressions that depend only on aggregate statistics of the labels and not on any individual label. Finally, we demonstrate how this framework can be used effectively to validate machine learning models that we trained on two real-world tasks from medical imaging. The discrepancy ratio metric reveals what conventional metrics do not: that our models not only vastly exceed the average human performance, but even exceed the performance of the best human experts in our datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 89fed659-d2ce-4c8c-b125-a7e444366322Related papers
- Forest vs Tree: The (N, K) Trade-off in Reproducible ML EvaluationDeepak Pandita, Flip Korn, Chris Welty, Christopher M. HomanAAAI 2026 · 2 citations
- Learning Calibrated Medical Image Segmentation via Multi-Rater Agreement ModelingWei Ji, Shuang Yu, Junde Wu, Kai Ma et al.CVPR 2021
- A Simple and Provable Approach for Learning on Noisy Labeled Medical ImagesNan Wang, Zonglin Di, Houlin He, Qingchao Jiang et al.ACM MM 2024 · 2 citations
- Disentangling Human Error from Ground Truth in Segmentation of Medical ImagesLe Zhang, Ryutaro Tanno, Moucheng Xu, Chen Jin et al.NeurIPS 2020 · 8 citations
- The Disagreement Deconvolution: Bringing Machine Learning Performance Metrics In Line With RealityMitchell L. Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto et al.CHI 2021 · 100 citations
