Beyond Bradley-Terry Models: A General Preference Model for Language Model Alignment
Yifan Zhang, Ge Zhang, Yue Wu, Kangping Xu, Quanquan Gu
摘要
Modeling human preferences is crucial for aligning foundation models with human values. Traditional reward modeling methods, such as the Bradley-Terry (BT) reward model, fall short in expressiveness, particularly in addressing intransitive preferences. In this paper, we introduce preference embedding, an approach that embeds responses into a latent space to capture intricate preference structures efficiently, achieving linear query complexity. Additionally, we propose preference score-based General Preference Optimization (GPO), which generalizes rewardbased reinforcement learning from human feedback (RLHF). Experimental results show that our General Preference embedding Model (GPM) consistently outperforms the BT reward model on the RewardBench benchmark and effectively models cyclic preferences where any BT reward model behaves like a random guess. Furthermore, evaluations on downstream tasks such as Al-pacaEval2.0, following the language model posttraining with GPO and our general preference model, reveal performance improvements over BT models. These findings indicate that our method may enhance the alignment of foundation models with nuanced human values. The code is available at https://github.com/general-preference/generalpreference-model .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- On the Design of KL-Regularized Policy Gradient Algorithms for LLM ReasoningYifan Zhang, Yifeng Liu, Rina Hughes, Yang Yuan 等ICLR 2026 · 被引用 30 次
- Ask a Strong LLM Judge when Your Reward Model is UncertainZhenghao Xu, Qin Lu, Qingru Zhang, Liang Qiu 等NeurIPS 2025 · 被引用 14 次
- Robust AI Evaluation through Maximal LotteriesHadi Khalaf, Serena Wang, Daniel Halpern, Itai Shapira 等ICML 2026 · 被引用 2 次
- What Does Preference Learning Recover from Pairwise Comparison Data?Rattana Pukdee, Nina Balcan, Pradeep RavikumarICML 2026 · 被引用 1 次
- Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-SolvingQi Liu, Xinhao Zheng, Renqiu Xia, Xingzhi Qi 等ICML 2026
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
相关 Paper
- Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret LearningYuheng Zhang, Dian Yu, Baolin Peng, Linfeng Song 等ICLR 2025
- Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model AlignmentYucong Huang, Xiucheng Li, Kaiqi Zhao, Jing LiICML 2026
- Permutative Preference Alignment from Listwise Ranking of Human JudgmentsYang Zhao, Yixin Wang, Mingzhang YinEMNLP 2025 · 被引用 5 次
- Greedy Sampling Is Provably Efficient For RLHFDi Wu, Chengshuai Shi, Jing Yang, Cong ShenNeurIPS 2025 · 被引用 11 次
- IPO: Your Language Model is Secretly a Preference ClassifierShivank Garg, Ayush Singh, Shweta Singh, Paras ChopraACL 2025
