Discovering Evolution Strategies via Meta-Black-Box Optimization
Robert Tjarko Lange, Tom Schaul, Yutian Chen, Tom Zahavy, Valentin Dalibard, Chris Lu, Satinder Singh, Sebastian Flennerhag
Abstract
Optimizing functions without access to gradients is the remit of black-box methods such as evolution strategies. While highly general, their learning dynamics are often times heuristic and inflexible --- exactly the limitations that meta-learning can address. Hence, we propose to discover effective update rules for evolution strategies via meta-learning. Concretely, our approach employs a search strategy parametrized by a self-attention-based architecture, which guarantees the update rule is invariant to the ordering of the candidate solutions. We show that meta-evolving this system on a small set of representative low-dimensional analytic optimization problems is sufficient to discover new evolution strategies capable of generalizing to unseen optimization problems, population sizes and optimization horizons. Furthermore, the same learned evolution strategy can outperform established neuroevolution baselines on supervised and continuous control tasks. As additional contributions, we ablate the individual neural network components of our method; reverse engineer the learned strategy into an explicit heuristic form, which remains highly competitive; and show that it is possible to self-referentially train an evolution strategy from scratch, with the learned update rule used to drive the outer meta-learning loop.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 83c5119f-dc0d-4253-af9c-114ca28b0422Cited by top-tier papers20
- ShinkaEvolve: Towards Open-Ended and Sample-Efficient Program EvolutionRobert T. Lange, Yuki Imajuku, Edoardo CetinICLR 2026 · 162 citations
- Darwin Gödel Machine: Open-Ended Evolution of Self-Improving AgentsJenny Zhang, Shengran Hu, Cong Lu, Robert Tjarko Lange et al.ICLR 2026 · 101 citations
- Discovering Preference Optimization Algorithms with and for Large Language ModelsChris Lu, Samuel Holt, Claudio Fanconi, Alex J. Chan et al.NeurIPS 2024 · 41 citations
- SYMBOL: Generating Flexible Black-Box Optimizers through Symbolic Equation LearningJiacheng Chen, Zeyuan Ma, Hongshu Guo, Yining Ma et al.ICLR 2024 · 27 citations
- B2Opt: Learning to Optimize Black-box Optimization with Little BudgetXiaobin Li, Kai Wu, Xiaoyu Zhang, Handing WangAAAI 2025 · 23 citations
Builds on13
- Perceiver: General Perception with Iterative AttentionAndrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals et al.ICML 2021 · 1,399 citations
- Discovering Reinforcement Learning AlgorithmsJunhyuk Oh, Matteo Hessel, Wojciech M. Czarnecki, Zhongwen Xu et al.NeurIPS 2020 · 154 citations
- Discovered Policy OptimisationChris Lu, Jakub Grudzien Kuba, Alistair Letcher, Luke Metz et al.NeurIPS 2022 · 134 citations
- Improving Generalization in Meta Reinforcement Learning using Learned ObjectivesLouis Kirsch, Sjoerd van Steenkiste, Jürgen SchmidhuberICLR 2020 · 132 citations
- A Self-Tuning Actor-Critic AlgorithmTom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel et al.NeurIPS 2020 · 106 citations
Related papers
- Discovering Temporally-Aware Reinforcement Learning AlgorithmsMatthew Thomas Jackson, Chris Lu, Louis Kirsch, Robert Tjarko Lange et al.ICLR 2024 · 23 citations
- Task-free Adaptive Meta Black-box OptimizationChao Wang, Licheng Jiao, Lingling Li, Jiaxuan Zhao et al.ICLR 2026 · 4 citations
- Neural Exploratory Landscape Analysis for Meta-Black-Box-OptimizationZeyuan Ma, Jiacheng Chen, Hongshu Guo, Yue-Jiao GongICLR 2025
- Low-Variance Gradient Estimation in Unrolled Computation Graphs with ES-SinglePaul VicolICML 2023 · 8 citations
- Learning Discrete Structured Variational Auto-Encoder using Natural Evolution StrategiesAlon Berliner, Guy Rotman, Yossi Adi, Roi Reichart et al.ICLR 2022 · 5 citations
