Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
Xin Qiu, Yulu Gan, Conor Hayes, Qiyao Liang, Yinggan XU, Roberto Dailey, Elliot Meyerson, Babak Hodjat, Risto Miikkulainen
摘要
Fine-tuning large language models (LLMs) for downstream tasks is an essential stage of modern AI deployment. Reinforcement learning (RL) has emerged as the dominant fine-tuning paradigm, underpinning many state-of-the-art LLMs. In contrast, evolution strategies (ES) has largely been overlooked due to the widespread belief that it does not scale to modern model sizes. This paper overturns this assumption by demonstrating the first successful application of ES to full-parameter fine-tuning of LLMs at the billion-parameter scale, without dimensionality reduction. ES can indeed search over extremely high-dimensional parameter spaces and outperform established RL implementations across multiple axes, including improved tolerance to long-horizon and delayed rewards, robustness across diverse base LLMs, reduced susceptibility to reward hacking, and improved training stability. These findings suggest that ES is not merely a viable alternative to RL, but a fundamentally different and powerful backpropagation-free post-training paradigm that opens a new direction for LLM fine-tuning beyond current RL-based approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Neural Thickets: Diverse Task Experts Are Dense Around Pretrained WeightsYulu Gan, Phillip IsolaICML 2026 · 被引用 17 次
- Evolution Strategies at the HyperscaleBidipta Sarkar, Mattie Fellows, Juan Duque, Alistair Letcher 等ICML 2026 · 被引用 16 次
- Procedural Generation Of Algorithm Discovery Tasks in Machine LearningAlexander D. Goldie, Zilin Wang, Adrian Hayler, Deepak Nathani 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
相关 Paper
- EvoLM: In Search of Lost Training Dynamics for Language Model ReasoningZhenting Qi, Fan Nie, Alexandre Alahi, James Y. Zou 等NeurIPS 2025 · 被引用 3 次
- Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language ModelsYinlam Chow, Guy Tennenholtz, Izzeddin Gur, Vincent Zhuang 等ICLR 2025 · 被引用 1 次
- Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement LearningHao Ma, Tianyi Hu, Zhiqiang Pu, Boyin Liu 等NeurIPS 2024 · 被引用 54 次
- Offline RL by Reward-Weighted Fine-Tuning for Conversation OptimizationSubhojyoti Mukherjee, Viet Dac Lai, Raghavendra Addanki, Ryan Rossi 等NeurIPS 2025 · 被引用 12 次
- EvolveSearch: An Iterative Self-Evolving Search AgentDingchu Zhang, Yida Zhao, Jialong Wu, Liwen Zhang 等EMNLP 2025 · 被引用 1 次
