ComPO: Preference Alignment via Comparison Oracles
Peter Chen, Xi Chen, Wotao Yin, Tianyi Lin
Abstract
Direct alignment methods are increasingly used for aligning large language models (LLMs) with human preferences. However, these methods suffer from the issues of likelihood displacement, which can be driven by noisy preference pairs that induce similar likelihood for preferred and dispreferred responses. The contributions of this paper are two-fold. First, we propose a preference alignment method based on zeroth-order, comparison-based optimization via comparison oracles and provide convergence guarantees for its basic mechanism. Second, we improve our method using some heuristics and conduct the experiments to demonstrate the flexibility and compatibility of practical mechanisms in improving the performance of LLMs using noisy preference pairs. Evaluations are conducted across multiple base and instruction-tuned models (Mistral-7B, Llama-3-8B and Gemma-2-9B) with benchmarks (AlpacaEval 2, MT-Bench and Arena-Hard) 1 . Experimental results show the effectiveness of our method as an alternative to addressing the limitations of existing methods, not only likelihood displacement but verbosity. A highlight of our work is that we evidence the importance of designing specialized methods for preference pairs with distinct likelihood margin, which complements the recent findings in Razin et al. [73].
- We identify that likelihood displacement issue is exacerbated by the ineffective handling of noisy preference pairs in existing methods. We propose to mitigate this issue by developing a method based on a specialized comparison oracle to extract useful information from these pairs. We also provide a convergence guarantee for the basic scheme of our method under non-convex, smooth settings. 2. To ensure computational efficiency for large-scale model fine-tuning, we enhance our method with several techniques, including integration with DPO to handle clean and noisy preference data separately and approximating expensive steps in standard comparison-oracle-based optimization by efficiently restricting and clipping normalized gradients. 3. We conduct extensive experiments demonstrating the flexibility and effectiveness of our practical approach in improving LLM performance, particularly leveraging both clean and noisy preference data. Evaluations are undertaken across base and instruction-tuned models (Mistral-7B, Llama-3-8B, and Gemma-2-9B) using benchmarks (AlpacaEval 2, MT-Bench, and Arena-Hard). Experimental results validate our approach's effectiveness, which addresses limitations in current direct alignment techniques.
Recent works [32,27,66,4,93,60,65] have shown that verbosity can be mitigated by incorporating appropriate regularization into the objective, suggesting that modifying the objective could better capture alignment goals. Although our method is not specifically designed for addressing verbosity issue, it consistently improve length-controlled win rate (LC), indicating that our method helps reduce verbosity by possibly optimizing a more robust and alignment-faithful objective.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 230b3f96-4958-4afe-a600-1497da28bbc2Cited by top-tier papers4
- Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious RewardPeter Chen, Xiaopeng Li, Ziniu Li, Wotao Yin et al.ICLR 2026 · 28 citations
- Reward-free Alignment for Conflicting ObjectivesPeter Chen, Xiaopeng Li, Xi Chen, Tianyi LinICML 2026 · 8 citations
- Noisy Pairwise-Comparison Random Search for Smooth Nonconvex OptimizationTaha EL BAKKALI EL KADI, Rayane Bouftini, Richard Zhang, Omar SaadiICML 2026 · 1 citation
- Mediocrity is the key for LLM as a Judge Anchor SelectionShachar Don-Yehiya, Asaf Yehudai, Leshem Choshen, Omri AbendACL 2026 · 1 citation
Builds on48
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
Related papers
- Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt DistillationAiwei Liu, Haoping Bai, Zhiyun Lu, Xiang Kong et al.ACL 2024 · 4 citations
- TODO: Enhancing LLM Alignment with Ternary PreferencesYuxiang Guo, Lu Yin, Bo Jiang, Jiaqi ZhangICLR 2025
- Spread Preference Annotation: Direct Preference Judgment for Efficient LLM AlignmentDongyoung Kim, Kimin Lee, Jinwoo Shin, Jaehyung KimICLR 2025
- ORPO: Monolithic Preference Optimization without Reference ModelJiwoo Hong, Noah Lee, James ThorneEMNLP 2024 · 71 citations
- AlphaPO: Reward Shape Matters for LLM AlignmentAman Gupta, Shao Tang, Qingquan Song, Sirou Zhu et al.ICML 2025
