What Does Preference Learning Recover from Pairwise Comparison Data?
Rattana Pukdee, Nina Balcan, Pradeep Ravikumar
Abstract
Pairwise preference learning is central to machine learning, with recent applications in aligning language models with human preferences. A typical dataset consists of triplets , where response is preferred over response for context . The Bradley--Terry (BT) model is the predominant approach, modeling preference probabilities as a function of latent score differences. Standard practice assumes data follows this model and learns the latent scores accordingly. However, real data may violate this assumption, and it remains unclear what BT learning recovers in such cases. Starting from triplet comparison data, we formalize the preference information it encodes through the conditional preference distribution (CPRD). We give precise conditions for when BT is appropriate for modeling the CPRD, and identify factors governing sample efficiency---namely, margin and connectivity. Together, these results offer a data-centric foundation for understanding what preference learning actually recovers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d5cd897-ab99-49eb-87e2-0018b2109292Builds on21
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
Related papers
- Reward Modeling with Ordinal Feedback: Wisdom of the CrowdShang Liu, Yu Pan, Guanting Chen, Xiaocheng LiICML 2025
- Preference Learning of Latent Decision Utilities with a Human-like Model of Preferential ChoiceSebastiaan De Peuter, Shibei Zhu, Yujia Guo, Andrew Howes et al.NeurIPS 2024 · 6 citations
- Generalizing while preserving monotonicity in comparison-based preference learning modelsJulien Fageot, Peva Blanchard, Gilles Bareilles, Lê-Nguyên HoangNeurIPS 2025 · 2 citations
- Axioms for Learning from Pairwise ComparisonsRitesh Noothigattu, Dominik Peters, Ariel D. ProcacciaNeurIPS 2020 · 23 citations
- Rethinking Reward Modeling in Preference-based Large Language Model AlignmentHao Sun, Yunyi Shen, Jean-Francois TonICLR 2025
