Reward Alignment Optimization: A Direct Point-wise Alignment Approach
Zelin Li, Jia Leng, Dawei Song, Yangen Hu
Abstract
Direct Alignment Algorithms (DAAs) such as DPO simplify RLHF by optimizing policies directly from preference pairs. However, the Bradley-Terry probability-gap objective can induce likelihood displacement and, under weak KL constraints, may even reduce the probability of preferred responses, while implicit rewards can be limited in generalization. We propose Reward Alignment Optimization (RAO), a point-wise direct alignment method that uses an explicit reward model to specify exact target generation probabilities and align the policy offline towards them. Our key insight is a theoretical principle we call "prefix consistency", which links the normalization terms of prompts that share a prefix. Leveraging this property, RAO decouples target reward differentials from bias terms, prevents decreasing preferred-response probabilities, and better exploits reward information both within and across prompts. Extensive experiments on multiple base LLMs show that RAO consistently outperforms existing DAAs while enabling controllable target probability distributions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2cff858-4707-4da6-95cc-869eb32974ccBuilds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- SimPO: Simple Preference Optimization with a Reference-Free RewardYu Meng, Mengzhou Xia, Danqi ChenNeurIPS 2024 · 1,203 citations
Related papers
- TokenRatio: Principled Token-Level Preference Optimization via Ratio MatchingTruong Nguyen, Tien-Phat Nguyen, Linh Van, Duy Nguyen et al.ICML 2026
- Displacement-Resistant Extensions of DPO with Nonconvex -DivergencesIdan Pipano, Shoham Sabach, Kavosh Asadi, Mohammad GhavamzadehICLR 2026
- The Differences Between Direct Alignment Algorithms are a BlurAlexey Gorbatovski, Boris Shaposhnikov, Viacheslav Sinii, Alexey Malakhov et al.ICML 2026
- AlphaPO: Reward Shape Matters for LLM AlignmentAman Gupta, Shao Tang, Qingquan Song, Sirou Zhu et al.ICML 2025
- Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language ModelsRei Higuchi, Taiji SuzukiICML 2025
