Calibrated Preference Learning: The Case of Label Ranking
Santo Thies, Viktor Bengs, Timo Kaufmann, Sebastian Vollmer, Eyke Hüllermeier
摘要
Calibration, the alignment of predicted probabilities with true outcome frequencies, is essential for reliable decision-making. While extensively studied for classification and regression, calibration has not been formally addressed for probabilistic label ranking, where the goal is to predict a distribution over orderings of a label set. Naively treating rankings as classes ignores their structure and fails to capture important modalities such as pairwise and top-k predictions. We formalize calibration for label ranking and develop a hierarchy of notions covering full rankings, sub-rankings, and top-k rankings. We prove that full-rank calibration implies the others but not conversely, and sub-ranking and top-k calibration are incomparable. Empirically, we find popular label ranking models are often poorly calibrated, with substantial differences between sub-ranking and top-k metrics. Applying our framework to RLHF reward models, we find that calibration correlates strongly but not perfectly with benchmark accuracy, suggesting it captures a meaningful quality dimension beyond top-1 accuracy. These findings motivate future work on understanding the downstream effects of miscalibration and developing methods to correct it.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz 等NeurIPS 2020 · 被引用 674 次
- RewardBench 2: Advancing Reward Model EvaluationSaumya Malik, Valentina Pyatkin, Sander Land, Jacob Morrison 等ICLR 2026 · 被引用 139 次
- A Consistent and Differentiable Lp Canonical Calibration Error EstimatorTeodora Popordanoska, Raphael Sayer, Matthew B. BlaschkoNeurIPS 2022 · 被引用 58 次
- Calibration tests beyond classificationDavid Widmann, Fredrik Lindsten, Dave ZachariahICLR 2021 · 被引用 23 次
相关 Paper
- Obtaining Calibrated Probabilities with Personalized Ranking ModelsWonbin Kweon, SeongKu Kang, Hwanjo YuAAAI 2022 · 被引用 20 次
- Meta-Cal: Well-controlled Post-hoc Calibration by RankingXingchen Ma, Matthew B. BlaschkoICML 2021 · 被引用 44 次
- Top-label calibration and multiclass-to-binary reductionsChirag Gupta, Aaditya RamdasICLR 2022 · 被引用 51 次
- Intra Order-preserving Functions for Calibration of Multi-Class Neural NetworksAmir Rahimi, Amirreza Shaban, Ching-An Cheng, Richard Hartley 等NeurIPS 2020 · 被引用 96 次
- Calibration by Distribution Matching: Trainable Kernel Calibration MetricsCharlie Marx, Sofian Zalouk, Stefano ErmonNeurIPS 2023 · 被引用 21 次
