Calibrated Preference Learning: The Case of Label Ranking
Santo Thies, Viktor Bengs, Timo Kaufmann, Sebastian Vollmer, Eyke Hüllermeier
Abstract
Calibration, the alignment of predicted probabilities with true outcome frequencies, is essential for reliable decision-making. While extensively studied for classification and regression, calibration has not been formally addressed for probabilistic label ranking, where the goal is to predict a distribution over orderings of a label set. Naively treating rankings as classes ignores their structure and fails to capture important modalities such as pairwise and top-k predictions. We formalize calibration for label ranking and develop a hierarchy of notions covering full rankings, sub-rankings, and top-k rankings. We prove that full-rank calibration implies the others but not conversely, and sub-ranking and top-k calibration are incomparable. Empirically, we find popular label ranking models are often poorly calibrated, with substantial differences between sub-ranking and top-k metrics. Applying our framework to RLHF reward models, we find that calibration correlates strongly but not perfectly with benchmark accuracy, suggesting it captures a meaningful quality dimension beyond top-1 accuracy. These findings motivate future work on understanding the downstream effects of miscalibration and developing methods to correct it.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3bda94fc-4a44-4b12-9d18-b8b5d7a94c41Builds on9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz et al.NeurIPS 2020 · 674 citations
- RewardBench 2: Advancing Reward Model EvaluationSaumya Malik, Valentina Pyatkin, Sander Land, Jacob Morrison et al.ICLR 2026 · 139 citations
- A Consistent and Differentiable Lp Canonical Calibration Error EstimatorTeodora Popordanoska, Raphael Sayer, Matthew B. BlaschkoNeurIPS 2022 · 58 citations
- Calibration tests beyond classificationDavid Widmann, Fredrik Lindsten, Dave ZachariahICLR 2021 · 23 citations
Related papers
- Obtaining Calibrated Probabilities with Personalized Ranking ModelsWonbin Kweon, SeongKu Kang, Hwanjo YuAAAI 2022 · 20 citations
- Meta-Cal: Well-controlled Post-hoc Calibration by RankingXingchen Ma, Matthew B. BlaschkoICML 2021 · 44 citations
- Top-label calibration and multiclass-to-binary reductionsChirag Gupta, Aaditya RamdasICLR 2022 · 51 citations
- Intra Order-preserving Functions for Calibration of Multi-Class Neural NetworksAmir Rahimi, Amirreza Shaban, Ching-An Cheng, Richard Hartley et al.NeurIPS 2020 · 96 citations
- Calibration by Distribution Matching: Trainable Kernel Calibration MetricsCharlie Marx, Sofian Zalouk, Stefano ErmonNeurIPS 2023 · 21 citations
