AtC: Aggregate-then-Calibrate for Human-centered Assessment
Zejun Xie, Xintong Li, Guang Wang, Desheng Zhang
摘要
Human-centered assessment tasks, which are essential for systematic decision-making, rely heavily on human judgment and typically lack verifiable ground truth. Existing approaches face a dilemma: methods using only human judgments suffer from heterogeneous expertise and inconsistent rating scales, while methods using only model-generated scores must learn from imperfect proxies or incomplete features. We propose Aggregate-then-Calibrate (AtC), a two-stage framework that combines these complementary sources. Stage-1 aggregates heterogeneous comparative judgments into a consensus ranking using a rank-aggregation model that accounts for annotator reliability. Stage-2 calibrates any predictive model’s scores by an isotonic projection onto the order , enforcing ordinal consistency while preserving as much of the model’s quantitative information as possible. Theoretically, we show: (1) modeling annotator heterogeneity yields strictly more efficient consensus estimation than homogeneity; (2) isotonic calibration enjoys risk bounds even when the consensus ranking is misspecified; and (3) AtC asymptotically outperforms model-only assessment. Across semi-synthetic and real-world datasets, AtC consistently improves accuracy and robustness over human-only or model-only assessments. Our results bridge judgment aggregation with model-free calibration, providing a principled recipe for human-centered assessment when ground truth is costly, scarce, or unverifiable.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- Rank Aggregation Algorithms for Fair ConsensusCaitlin Kuhlman, Elke A. RundensteinerVLDB 2020 · 被引用 60 次
- You Are the Best Reviewer of Your Own Papers: An Owner-Assisted Scoring MechanismWeijie J. SuNeurIPS 2021 · 被引用 32 次
- Rank Aggregation via Heterogeneous Thurstone Preference ModelsTao Jin, Pan Xu, Quanquan Gu, Farzad FarnoudAAAI 2020 · 被引用 19 次
- Human Expertise in Algorithmic PredictionRohan Alur, Manish Raghavan, Devavrat ShahNeurIPS 2024 · 被引用 18 次
- Explaining Preferences with Shapley ValuesRobert Hu, Siu Lun Chau, Jaime Ferrando Huertas, Dino SejdinovicNeurIPS 2022 · 被引用 11 次
相关 Paper
- STABLEVAL: Disagreement-Aware and Stable Evaluation of AI SystemsSailendra Akash Bonagiri, Gerard Anderias, Saee Patil, Angelina Lai 等ICML 2026 · 被引用 2 次
- Causal Isotonic Calibration for Heterogeneous Treatment EffectsLars van der Laan, Ernesto Ulloa-Pérez, Marco Carone, Alex LuedtkeICML 2023 · 被引用 18 次
- Uncalibrated Models Can Improve Human-AI CollaborationKailas Vodrahalli, Tobias Gerstenberg, James Y. ZouNeurIPS 2022 · 被引用 47 次
- A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground TruthMingyuan Xu, Xinzi Tan, Jiawei Wu, Doudou ZhouICML 2026
- Epistemic Uncertainty Quantification To Improve Decisions From Black-Box ModelsSébastien Melo, Gaël Varoquaux, Marine Le MorvanICLR 2026
