On-the-fly Preference Alignment via Principle-Guided Decoding
Mingye Zhu, Yi Liu, Lei Zhang, Junbo Guo, Zhendong Mao
Abstract
With the rapidly expanding landscape of large language models, aligning model generations with human values and preferences is becoming increasingly important. Popular alignment methods, such as Reinforcement Learning from Human Feedback, have shown significant success in guiding models with greater control. However, these methods require considerable computational resources, which is inefficient, and substantial collection of training data to accommodate the diverse and pluralistic nature of human preferences, which is impractical. These limitations significantly constrain the scope and efficacy of both task-specific and general preference alignment methods. In this work, we introduce On-the-fly Preference Alignment via Principle-Guided Decoding (OPAD) to directly align model outputs with human preferences during inference, eliminating the need for fine-tuning. Our approach involves first curating a surrogate solution to an otherwise infeasible optimization problem and then designing a principle-guided reward function based on this surrogate. The final aligned policy is derived by maximizing this customized reward, which exploits the discrepancy between the constrained policy and its unconstrained counterpart. OPAD directly modifies the model's predictions during inference, ensuring principle adherence without incurring the computational overhead of retraining or fine-tuning. Experiments show that OPAD achieves competitive or superior performance in both general and personalized alignment tasks, demonstrating its efficiency and effectiveness compared to state-of-the-art baselines. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a84ffd2-3049-44c9-9c59-4a25b6ebbad1Cited by top-tier papers6
- Inference-Time Personalized Alignment with a Few User Preference QueriesVictor-Alexandru Padurean, Parameswaran Kamalaruban, Nachiket Kotalwar, Alkis Gotovos et al.NeurIPS 2025 · 5 citations
- Safety Alignment of Large Language Models via Contrasting Safe and Harmful DistributionsXiaoyun Zhang, Zhengyue Zhao, Wenxuan Shi, Kaidi Xu et al.AAAI 2026 · 4 citations
- LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference OptimizationJunsong Li, Jie Zhou, Bihao Zhan, Yutao Yang et al.AAAI 2026 · 3 citations
- Enhancing Persona Following at Decoding Time via Dynamic Importance Estimation for Role-Playing AgentsYuxin Liu, Mingye Zhu, Siyuan Liu, Bo Hu et al.ICLR 2026 · 2 citations
- Leveraging Importance Sampling to Detach Alignment Modules from Large Language ModelsYi Liu, Dianqing Liu, Mingye Zhu, Junbo Guo et al.NeurIPS 2025 · 1 citation
Builds on12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- The Unlocking Spell on Base LLMs: Rethinking Alignment via In-Context LearningBill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri et al.ICLR 2024 · 299 citations
- RAIN: Your Language Models Can Align Themselves without FinetuningYuhui Li, Fangyun Wei, Jinjing Zhao, Chao Zhang et al.ICLR 2024 · 171 citations
- Residual Energy-Based Models for Text GenerationYuntian Deng, Anton Bakhtin, Myle Ott, Arthur Szlam et al.ICLR 2020 · 147 citations
Related papers
- Alignment-Aware DecodingFrédéric Berdoz, Luca Lanzendörfer, René Caky, Roger WattenhoferICML 2026 · 1 citation
- PAD: Personalized Alignment of LLMs at Decoding-timeRuizhe Chen, Xiaotian Zhang, Meng Luo, Wenhao Chai et al.ICLR 2025
- ActiveDPO: Active Direct Preference Optimization for Sample-Efficient AlignmentXiaoqiang Lin, Arun Verma, Zhongxiang Dai, Daniela Rus et al.ICLR 2026 · 12 citations
- Self-Guided Alignment: Adaptive Preference Sensing for Multi-Objective GenerationNing Wang, Zhanyang Liu, Taotao Zhou, Xinrui Zhang et al.ACL 2026
- Fast Best-of-N Decoding via Speculative RejectionHanshi Sun, Momin Haider, Ruiqi Zhang, Huitao Yang et al.NeurIPS 2024 · 144 citations
