Decoding-Time Language Model Alignment with Multiple Objectives
Ruizhe Shi, Yifang Chen, Yushi Hu, Alisa Liu, Hanna Hajishirzi, Noah A. Smith, Simon S. Du
Abstract
Aligning language models (LMs) to human preferences has emerged as a critical pursuit, enabling these models to better serve diverse user needs. Existing methods primarily focus on optimizing LMs for a single reward function, limiting their adaptability to varied objectives. Here, we propose , a decoding-time algorithm that outputs the next token from a linear combination of predictions of all base models, for any given weightings over different objectives. We exploit a common form among a family of -divergence regularized alignment approaches (such as PPO, DPO, and their variants) to identify a closed-form solution by Legendre transform, and derive an efficient decoding strategy. Theoretically, we show why existing approaches can be sub-optimal even in natural settings and obtain optimality guarantees for our method. Empirical results demonstrate the effectiveness of the algorithm. For example, compared to a parameter-merging baseline, MOD achieves 12.8% overall reward improvement when equally optimizing towards objectives. Moreover, we experiment with MOD on combining three fully-finetuned LLMs of different model sizes, each aimed at different objectives such as safety, coding, and general user preference. Unlike traditional methods that require careful curation of a mixture of datasets to achieve comprehensive improvement, we can quickly experiment with preference weightings using MOD to find the best combination of models. Our best combination reduces toxicity on Toxigen to nearly 0% and achieves 7.9--33.3% improvement across other three metrics (, Codex@1, GSM-COT, BBH-COT).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 78dfe10f-2222-4b0e-9809-2d0d297a1f47Cited by top-tier papers47
- Multi-Attribute Steering of Language Models via Targeted InterventionDuy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit BansalACL 2025 · 30 citations
- When One LLM Drools, Multi-LLM Collaboration RulesShangbin Feng, Wenxuan Ding, Alisa Liu, Zifeng Wang et al.ACL 2026 · 27 citations
- Mechanism Design for LLM Fine-tuning with Multiple Reward ModelsHaoran Sun, Yurong Chen, Siwei Wang, Chu Xu et al.NeurIPS 2025 · 26 citations
- Representation-Based Exploration for Language Models: From Test-Time to Post-TrainingJens Tuyls, Dylan J Foster, Akshay Krishnamurthy, Jordan T. AshICLR 2026 · 18 citations
- Multi-objective Large Language Model Alignment with Hierarchical ExpertsZhuo Li, Guodong DU, Weiyang Guo, Yigeng Zhou et al.ICLR 2026 · 17 citations
Builds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
Related papers
- Robust Multi-Objective Controlled Decoding of Large Language ModelsSeongho Son, William Bankes, Sangwoong Yoon, Shyam Sundhar Ramesh et al.ICLR 2026 · 12 citations
- Collab: Controlled Decoding using Mixture of Agents for LLM AlignmentSouradip Chakraborty, Sujay Bhatt, Udari Madhushani Sehwag, Soumya Suvra Ghosal et al.ICLR 2025
- Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable RewardsYiran Shen, Yu Xia, Jonathan Chang, Prithviraj AmmanabroluICML 2026
- Towards Aligning Language Models with Textual FeedbackSaüc Abadal Lloret, Shehzaad Dhuliawala, Keerthiram Murugesan, Mrinmaya SachanEMNLP 2024 · 2 citations
- Controlled Decoding from Language ModelsSidharth Mudgal, Jong Lee, Harish Ganapathy, YaGuang Li et al.ICML 2024 · 130 citations
