STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems
Sailendra Akash Bonagiri, Gerard Anderias, Saee Patil, Angelina Lai, Devang Borkar, Gezheng Kang, Ishant Gandhi, Setareh Rafatirad, Houman Homayoun
Abstract
Human evaluation remains the primary standard for assessing modern AI systems, yet annotator disagreement, bias, and variability make system rankings fragile under standard majority vote aggregation. Majority vote discards annotator reliability and item-level ambiguity, often yielding unstable comparisons across annotator subsets. We introduce STABLEVAL, a disagreement-aware evaluation framework that models latent item correctness and annotator-specific confusion patterns to produce posterior expected item credit and calibrated agent-level scores. Unlike label-denoising approaches such as Dawid--Skene, STABLEVAL is explicitly designed for stable and uncertainty-aware system evaluation rather than hard label recovery. We formalize ranking stability as a first-class evaluation objective and analyze how aggregation methods preserve or distort underlying annotator behavior. Across controlled synthetic experiments and multiple real-world human-annotated benchmarks, majority vote exhibits increasing score error and ranking instability under annotator heterogeneity and adversarial noise, while STABLEVAL yields more stable and statistically grounded system rankings. These results demonstrate that modeling disagreement is essential for robust and reproducible AI evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef3725c1-566a-4cd6-82e5-037b3cf11e31Builds on5
- Asking and Answering Questions to Evaluate the Factual Consistency of SummariesAlex Wang, Kyunghyun Cho, Mike LewisACL 2020 · 317 citations
- When the Majority is Wrong: Modeling Annotator Disagreement for Subjective TasksEve Fleisig, Rediet Abebe, Dan KleinEMNLP 2023 · 11 citations
- Modeling Annotator Disagreement with Demographic-Aware Experts and Synthetic PerspectivesYinuo Xu, Veronica Derricks, Allison Earl, David JurgensACL 2026 · 8 citations
- The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMsNitay Calderon, Roi Reichart, Rotem DrorACL 2025
- Bayesian Calibration of Win Rate Estimation with LLM EvaluatorsYicheng Gao, Gonghan Xu, Zhe Wang, Arman CohanEMNLP 2024
Related papers
- Dependence-Aware Label Aggregation for LLM-as-a-Judge via Ising ModelsKrishna Balasubramanian, Aleksandr Podkopaev, Shiva KasiviswanathanICML 2026
- Learning Calibrated Medical Image Segmentation via Multi-Rater Agreement ModelingWei Ji, Shuang Yu, Junde Wu, Kai Ma et al.CVPR 2021
- AtC: Aggregate-then-Calibrate for Human-centered AssessmentZejun Xie, Xintong Li, Guang Wang, Desheng ZhangICLR 2026
- Discrepancy Ratio: Evaluating Model Performance When Even Experts Disagree on the TruthIgor Lovchinsky, Alon Daks, Israel Malkin, Pouya Samangouei et al.ICLR 2020 · 11 citations
- PERSEVAL: A Framework for Perspectivist Classification EvaluationSoda Marem Lo, Silvia Casola, Erhan Sezerer, Valerio Basile et al.EMNLP 2025
