Transfer Q-star : Principled Decoding for LLM Alignment
Souradip Chakraborty, Soumya Suvra Ghosal, Ming Yin, Dinesh Manocha, Mengdi Wang, Amrit Singh Bedi, Furong Huang
Abstract
Aligning foundation models is essential for their safe and trustworthy deployment. However, traditional fine-tuning methods are computationally intensive and require updating billions of model parameters. A promising alternative, alignment via decoding, adjusts the response distribution directly without model updates to maximize a target reward , thus providing a lightweight and adaptable framework for alignment. However, principled decoding methods rely on oracle access to an optimal Q-function (), which is often unavailable in practice. Hence, prior SoTA methods either approximate this using (derived from the reference model) or rely on short-term rewards, resulting in sub-optimal decoding performance. In this work, we propose Transfer , which implicitly estimates the optimal value function for a target reward through a baseline model aligned with a baseline reward (which can be different from the target reward ). Theoretical analyses of Transfer provide a rigorous characterization of its optimality, deriving an upper bound on the sub-optimality gap and identifying a hyperparameter to control the deviation from the pre-trained reference model based on user needs. Our approach significantly reduces the sub-optimality gap observed in prior SoTA methods and demonstrates superior empirical performance across key metrics such as coherence, diversity, and quality in extensive tests on several synthetic and real datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da4cb8a4-11ec-41be-a658-17eaf12ebc29Cited by top-tier papers23
- Taming Imperfect Process Verifiers: A Sampling Perspective on BacktrackingDhruv Rohatgi, Abhishek Shetty, Donya Saless, Yuchen Li et al.ICLR 2026 · 15 citations
- Test-Time Scaling in Diffusion LLMS via Hidden Semi-Autoregressive ExpertsJihoon Lee, Hoyeon Moon, Kevin Zhai, Arun Kumar Chithanar et al.ICLR 2026 · 7 citations
- Leveraging Machine Unlearning for Cost-Efficient Preference AlignmentXiaoHua Feng, Yuyuan Li, HuWei Ji, Li Zhang et al.ICML 2026 · 4 citations
- From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time AlignmentBin Xie, Bingbing Xu, Yige Yuan, Shengmao Zhu et al.ACL 2025 · 4 citations
- Constrained Auto-Regressive Decoding Constrains Generative RetrievalShiguang Wu, Zhaochun Ren, Xin Xin, Jiyuan Yang et al.SIGIR 2025 · 3 citations
Builds on17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
Related papers
- InfAlign: Inference-aware language model alignmentAnanth Balashankar, Ziteng Sun, Jonathan Berant, Jacob Eisenstein et al.ICML 2025
- Joint Reward and Policy Learning with Demonstrations and Human Feedback Improves AlignmentChenliang Li, Siliang Zeng, Zeyi Liao, Jiaxiang Li et al.ICLR 2025
- Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference AdjustmentRui Yang, Xiaoman Pan, Feng Luo, Shuang Qiu et al.ICML 2024 · 144 citations
- DARC: Disagreement-Aware Alignment via Risk-Constrained Decodingmingxi Zou, Jiaxiang Chen, Junfan Li, Langzhang Liang et al.ICML 2026
- Moral Alignment for LLM AgentsElizaveta Tennant, Stephen Hailes, Mirco MusolesiICLR 2025
