Lune

NeurIPS2025Top-tier venue

ComPO: Preference Alignment via Comparison Oracles

Peter Chen, Xi Chen, Wotao Yin, Tianyi Lin

2025Year
20Citations
4Top-tier citations

Abstract

Direct alignment methods are increasingly used for aligning large language models (LLMs) with human preferences. However, these methods suffer from the issues of likelihood displacement, which can be driven by noisy preference pairs that induce similar likelihood for preferred and dispreferred responses. The contributions of this paper are two-fold. First, we propose a preference alignment method based on zeroth-order, comparison-based optimization via comparison oracles and provide convergence guarantees for its basic mechanism. Second, we improve our method using some heuristics and conduct the experiments to demonstrate the flexibility and compatibility of practical mechanisms in improving the performance of LLMs using noisy preference pairs. Evaluations are conducted across multiple base and instruction-tuned models (Mistral-7B, Llama-3-8B and Gemma-2-9B) with benchmarks (AlpacaEval 2, MT-Bench and Arena-Hard) 1 . Experimental results show the effectiveness of our method as an alternative to addressing the limitations of existing methods, not only likelihood displacement but verbosity. A highlight of our work is that we evidence the importance of designing specialized methods for preference pairs with distinct likelihood margin, which complements the recent findings in Razin et al. [73].

  1. We identify that likelihood displacement issue is exacerbated by the ineffective handling of noisy preference pairs in existing methods. We propose to mitigate this issue by developing a method based on a specialized comparison oracle to extract useful information from these pairs. We also provide a convergence guarantee for the basic scheme of our method under non-convex, smooth settings. 2. To ensure computational efficiency for large-scale model fine-tuning, we enhance our method with several techniques, including integration with DPO to handle clean and noisy preference data separately and approximating expensive steps in standard comparison-oracle-based optimization by efficiently restricting and clipping normalized gradients. 3. We conduct extensive experiments demonstrating the flexibility and effectiveness of our practical approach in improving LLM performance, particularly leveraging both clean and noisy preference data. Evaluations are undertaken across base and instruction-tuned models (Mistral-7B, Llama-3-8B, and Gemma-2-9B) using benchmarks (AlpacaEval 2, MT-Bench, and Arena-Hard). Experimental results validate our approach's effectiveness, which addresses limitations in current direct alignment techniques.

Recent works [32,27,66,4,93,60,65] have shown that verbosity can be mitigated by incorporating appropriate regularization into the objective, suggesting that modifying the objective could better capture alignment goals. Although our method is not specifically designed for addressing verbosity issue, it consistently improve length-controlled win rate (LC), indicating that our method helps reduce verbosity by possibly optimizing a more robust and alignment-faithful objective.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 230b3f96-4958-4afe-a600-1497da28bbc2

Cited by top-tier papers4

Ask how each one uses it

Builds on48

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines