Beyond Pairwise: Empowering LLM Alignment With (Ranked) Choice Modeling
Yuxuan Tang, Yifan Feng
Abstract
Alignment of large language models (LLMs) has predominantly relied on pairwise preference optimization, where annotators select the better of two responses to a prompt. While simple, this approach overlooks the opportunity to learn from richer forms of human feedback, such as multiway comparisons and top- rankings. We introduce Ranked Choice Preference Optimization (RCPO), a unified framework that bridges preference optimization with (ranked) choice modeling via maximum likelihood estimation. RCPO supports both utility-based and rank-based models, subsumes several pairwise methods (such as DPO and SimPO) as special cases, and provides principled training objectives for richer feedback formats. We instantiate this framework with two representative models (Multinomial Logit and Mallows-RMJ). Experiments on Llama-3-8B-Instruct, Gemma-2-9B-it, and Mistral-7B-Instruct across in-distribution and out-of-distribution settings show that RCPO consistently outperforms competitive baselines. RCPO shows that directly leveraging ranked preference data, combined with the right choice models, yields more effective alignment. It offers an extensible foundation for incorporating (ranked) choice modeling into LLM training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47beb673-1037-485a-9833-6225be5bc032Cited by top-tier papers1
Ask how each one uses itBuilds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky et al.ICML 2024 · 973 citations
Related papers
- Latent Preference Coding: Aligning Large Language Models via Discrete Latent CodesZhuocheng Gong, Jian Guan, Wei Wu, Huishuai Zhang et al.ICML 2025
- RE-PO: Robust Enhanced Policy Optimization as a General Framework for LLM AlignmentXiaoyang Cao, Zelai Xu, Mo Guang, Kaiwen Long et al.ICLR 2026 · 4 citations
- AlphaDPO: Adaptive Reward Margin for Direct Preference OptimizationJunkang Wu, Xue Wang, Zhengyi Yang, Jiancan Wu et al.ICML 2025
- AlphaPO: Reward Shape Matters for LLM AlignmentAman Gupta, Shao Tang, Qingquan Song, Sirou Zhu et al.ICML 2025
- RMO: Towards Better LLM Alignment via Reshaping Reward Margin DistributionsYanchi Ru, Yue Huang, Xiangliang ZhangAAAI 2026 · 1 citation
