AtC: Aggregate-then-Calibrate for Human-centered Assessment
Zejun Xie, Xintong Li, Guang Wang, Desheng Zhang
Abstract
Human-centered assessment tasks, which are essential for systematic decision-making, rely heavily on human judgment and typically lack verifiable ground truth. Existing approaches face a dilemma: methods using only human judgments suffer from heterogeneous expertise and inconsistent rating scales, while methods using only model-generated scores must learn from imperfect proxies or incomplete features. We propose Aggregate-then-Calibrate (AtC), a two-stage framework that combines these complementary sources. Stage-1 aggregates heterogeneous comparative judgments into a consensus ranking using a rank-aggregation model that accounts for annotator reliability. Stage-2 calibrates any predictive model’s scores by an isotonic projection onto the order , enforcing ordinal consistency while preserving as much of the model’s quantitative information as possible. Theoretically, we show: (1) modeling annotator heterogeneity yields strictly more efficient consensus estimation than homogeneity; (2) isotonic calibration enjoys risk bounds even when the consensus ranking is misspecified; and (3) AtC asymptotically outperforms model-only assessment. Across semi-synthetic and real-world datasets, AtC consistently improves accuracy and robustness over human-only or model-only assessments. Our results bridge judgment aggregation with model-free calibration, providing a principled recipe for human-centered assessment when ground truth is costly, scarce, or unverifiable.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9a22337-fb90-4b98-a127-cb62511aecd2Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Rank Aggregation Algorithms for Fair ConsensusCaitlin Kuhlman, Elke A. RundensteinerVLDB 2020 · 60 citations
- You Are the Best Reviewer of Your Own Papers: An Owner-Assisted Scoring MechanismWeijie J. SuNeurIPS 2021 · 32 citations
- Rank Aggregation via Heterogeneous Thurstone Preference ModelsTao Jin, Pan Xu, Quanquan Gu, Farzad FarnoudAAAI 2020 · 19 citations
- Human Expertise in Algorithmic PredictionRohan Alur, Manish Raghavan, Devavrat ShahNeurIPS 2024 · 18 citations
- Explaining Preferences with Shapley ValuesRobert Hu, Siu Lun Chau, Jaime Ferrando Huertas, Dino SejdinovicNeurIPS 2022 · 11 citations
Related papers
- STABLEVAL: Disagreement-Aware and Stable Evaluation of AI SystemsSailendra Akash Bonagiri, Gerard Anderias, Saee Patil, Angelina Lai et al.ICML 2026 · 2 citations
- Causal Isotonic Calibration for Heterogeneous Treatment EffectsLars van der Laan, Ernesto Ulloa-Pérez, Marco Carone, Alex LuedtkeICML 2023 · 18 citations
- Uncalibrated Models Can Improve Human-AI CollaborationKailas Vodrahalli, Tobias Gerstenberg, James Y. ZouNeurIPS 2022 · 47 citations
- A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground TruthMingyuan Xu, Xinzi Tan, Jiawei Wu, Doudou ZhouICML 2026
- Epistemic Uncertainty Quantification To Improve Decisions From Black-Box ModelsSébastien Melo, Gaël Varoquaux, Marine Le MorvanICLR 2026
