Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models
Zhanhui Zhou, Zhixuan Liu, Jie Liu, Zhichen Dong, Chao Yang, Yu Qiao
Abstract
Large language models are usually fine-tuned to align with human preferences. However, fine-tuning a large language model can be challenging. In this work, we introduce , framing the alignment of a large language model as a test-time greedy search to maximize the log-probability difference between small tuned and untuned models while sampling from the frozen large model. This method serves both as (1) a compute-efficient model up-scaling strategy that avoids directly tuning the large model and as (2) an instance of weak-to-strong generalization that enhances a strong model with weak test-time guidance. Empirically, we demonstrate the flexibility of weak-to-strong search across different tasks. In controlled-sentiment generation and summarization, we use tuned and untuned s to improve the alignment of large models without additional training. Crucially, in a more difficult instruction-following benchmark, AlpacaEval 2.0, we show that reusing off-the-shelf small models (e.g., and its untuned version) can improve the length-controlled win rates of both white-box and black-box large models against (e.g., for and for ), despite the small models' low win rates .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 094092a1-81ef-463e-aa76-ffd4704936d1Cited by top-tier papers19
- Discovering Preference Optimization Algorithms with and for Large Language ModelsChris Lu, Samuel Holt, Claudio Fanconi, Alex J. Chan et al.NeurIPS 2024 · 41 citations
- Robust Multi-Objective Controlled Decoding of Large Language ModelsSeongho Son, William Bankes, Sangwoong Yoon, Shyam Sundhar Ramesh et al.ICLR 2026 · 12 citations
- Inference-time Alignment in Continuous SpaceYige Yuan, Teng Xiao, Yunfan Li, Bingbing Xu et al.NeurIPS 2025 · 9 citations
- Leveraging Machine Unlearning for Cost-Efficient Preference AlignmentXiaoHua Feng, Yuyuan Li, HuWei Ji, Li Zhang et al.ICML 2026 · 4 citations
- From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time AlignmentBin Xie, Bingbing Xu, Yige Yuan, Shengmao Zhu et al.ACL 2025 · 4 citations
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky et al.ICML 2024 · 973 citations
Related papers
- Weak-to-Strong Preference Optimization: Stealing Reward from Weak Aligned ModelWenhong Zhu, Zhiwei He, Xiaofeng Wang, Pengfei Liu et al.ICLR 2025
- SimPER: A Minimalist Approach to Preference Alignment without HyperparametersTeng Xiao, Yige Yuan, Zhengyu Chen, Mingxiao Li et al.ICLR 2025
- Improving Model Alignment Through Collective Intelligence of Open-Source ModelsJunlin Wang, Roy Xie, Shang Zhu, Jue Wang et al.ICML 2025
- ORPO: Monolithic Preference Optimization without Reference ModelJiwoo Hong, Noah Lee, James ThorneEMNLP 2024 · 71 citations
- W2S-AlignTree: Weak-to-Strong Inference-Time Alignment for Large Language Models via Monte Carlo Tree SearchZhenyu Ding, Yuhao Wang, Tengyue Xiao, Haoying Wang et al.AAAI 2026
