Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models
Zhanhui Zhou, Zhixuan Liu, Jie Liu, Zhichen Dong, Chao Yang, Yu Qiao
摘要
Large language models are usually fine-tuned to align with human preferences. However, fine-tuning a large language model can be challenging. In this work, we introduce , framing the alignment of a large language model as a test-time greedy search to maximize the log-probability difference between small tuned and untuned models while sampling from the frozen large model. This method serves both as (1) a compute-efficient model up-scaling strategy that avoids directly tuning the large model and as (2) an instance of weak-to-strong generalization that enhances a strong model with weak test-time guidance. Empirically, we demonstrate the flexibility of weak-to-strong search across different tasks. In controlled-sentiment generation and summarization, we use tuned and untuned s to improve the alignment of large models without additional training. Crucially, in a more difficult instruction-following benchmark, AlpacaEval 2.0, we show that reusing off-the-shelf small models (e.g., and its untuned version) can improve the length-controlled win rates of both white-box and black-box large models against (e.g., for and for ), despite the small models' low win rates .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Discovering Preference Optimization Algorithms with and for Large Language ModelsChris Lu, Samuel Holt, Claudio Fanconi, Alex J. Chan 等NeurIPS 2024 · 被引用 41 次
- Robust Multi-Objective Controlled Decoding of Large Language ModelsSeongho Son, William Bankes, Sangwoong Yoon, Shyam Sundhar Ramesh 等ICLR 2026 · 被引用 12 次
- Inference-time Alignment in Continuous SpaceYige Yuan, Teng Xiao, Yunfan Li, Bingbing Xu 等NeurIPS 2025 · 被引用 9 次
- Leveraging Machine Unlearning for Cost-Efficient Preference AlignmentXiaoHua Feng, Yuyuan Li, HuWei Ji, Li Zhang 等ICML 2026 · 被引用 4 次
- From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time AlignmentBin Xie, Bingbing Xu, Yige Yuan, Shengmao Zhu 等ACL 2025 · 被引用 4 次
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky 等ICML 2024 · 被引用 973 次
相关 Paper
- Weak-to-Strong Preference Optimization: Stealing Reward from Weak Aligned ModelWenhong Zhu, Zhiwei He, Xiaofeng Wang, Pengfei Liu 等ICLR 2025
- SimPER: A Minimalist Approach to Preference Alignment without HyperparametersTeng Xiao, Yige Yuan, Zhengyu Chen, Mingxiao Li 等ICLR 2025
- Improving Model Alignment Through Collective Intelligence of Open-Source ModelsJunlin Wang, Roy Xie, Shang Zhu, Jue Wang 等ICML 2025
- ORPO: Monolithic Preference Optimization without Reference ModelJiwoo Hong, Noah Lee, James ThorneEMNLP 2024 · 被引用 71 次
- W2S-AlignTree: Weak-to-Strong Inference-Time Alignment for Large Language Models via Monte Carlo Tree SearchZhenyu Ding, Yuhao Wang, Tengyue Xiao, Haoying Wang 等AAAI 2026
