Generalizing while preserving monotonicity in comparison-based preference learning models
Julien Fageot, Peva Blanchard, Gilles Bareilles, Lê-Nguyên Hoang
Abstract
If you tell a learning model that you prefer an alternative over another alternative , then you probably expect the model to be monotone, that is, the valuation of increases, and that of decreases. Yet, perhaps surprisingly, many widely deployed comparison-based preference learning models, including large language models, fail to have this guarantee. Until now, the only comparison-based preference learning algorithms that were proved to be monotone are the Generalized Bradley-Terry models. Yet, these models are unable to generalize to uncompared data. In this paper, we advance the understanding of the set of models with generalization ability that are monotone. Namely, we propose a new class of Linear Generalized Bradley-Terry models with Diffusion Priors, and identify sufficient conditions on alternatives'embeddings that guarantee monotonicity. Our experiments show that this monotonicity is far from being a general guarantee, and that our new class of generalizing models improves accuracy, especially when the dataset is limited.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 51904662-b5d1-44d8-9b91-1bb0a2202e04Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Generalized Preference Optimization: A Unified Approach to Offline AlignmentYunhao Tang, Zhaohan Daniel Guo, Zeyu Zheng, Daniele Calandriello et al.ICML 2024 · 159 citations
- Learning to summarize with human feedbackNisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler et al.NeurIPS 2020 · 124 citations
- Preference Learning Algorithms Do Not Learn Preference RankingsAngelica Chen, Sadhika Malladi, Lily H. Zhang, Xinyi Chen et al.NeurIPS 2024 · 60 citations
Related papers
- Rethinking Reward Modeling in Preference-based Large Language Model AlignmentHao Sun, Yunyi Shen, Jean-Francois TonICLR 2025
- Greedy Sampling Is Provably Efficient For RLHFDi Wu, Chengshuai Shi, Jing Yang, Cong ShenNeurIPS 2025 · 11 citations
- Beyond Bradley-Terry Models: A General Preference Model for Language Model AlignmentYifan Zhang, Ge Zhang, Yue Wu, Kangping Xu et al.ICML 2025
- Generalized Bradley-Terry Models for Score Estimation from Paired ComparisonsJulien Fageot, Sadegh Farhadkhani, Lê-Nguyên Hoang, Oscar VillemaudAAAI 2024 · 12 citations
- What Does Preference Learning Recover from Pairwise Comparison Data?Rattana Pukdee, Nina Balcan, Pradeep RavikumarICML 2026 · 1 citation
