Discovering Evolution Strategies via Meta-Black-Box Optimization
Robert Tjarko Lange, Tom Schaul, Yutian Chen, Tom Zahavy, Valentin Dalibard, Chris Lu, Satinder Singh, Sebastian Flennerhag
摘要
Optimizing functions without access to gradients is the remit of black-box methods such as evolution strategies. While highly general, their learning dynamics are often times heuristic and inflexible --- exactly the limitations that meta-learning can address. Hence, we propose to discover effective update rules for evolution strategies via meta-learning. Concretely, our approach employs a search strategy parametrized by a self-attention-based architecture, which guarantees the update rule is invariant to the ordering of the candidate solutions. We show that meta-evolving this system on a small set of representative low-dimensional analytic optimization problems is sufficient to discover new evolution strategies capable of generalizing to unseen optimization problems, population sizes and optimization horizons. Furthermore, the same learned evolution strategy can outperform established neuroevolution baselines on supervised and continuous control tasks. As additional contributions, we ablate the individual neural network components of our method; reverse engineer the learned strategy into an explicit heuristic form, which remains highly competitive; and show that it is possible to self-referentially train an evolution strategy from scratch, with the learned update rule used to drive the outer meta-learning loop.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- ShinkaEvolve: Towards Open-Ended and Sample-Efficient Program EvolutionRobert T. Lange, Yuki Imajuku, Edoardo CetinICLR 2026 · 被引用 162 次
- Darwin Gödel Machine: Open-Ended Evolution of Self-Improving AgentsJenny Zhang, Shengran Hu, Cong Lu, Robert Tjarko Lange 等ICLR 2026 · 被引用 101 次
- Discovering Preference Optimization Algorithms with and for Large Language ModelsChris Lu, Samuel Holt, Claudio Fanconi, Alex J. Chan 等NeurIPS 2024 · 被引用 41 次
- SYMBOL: Generating Flexible Black-Box Optimizers through Symbolic Equation LearningJiacheng Chen, Zeyuan Ma, Hongshu Guo, Yining Ma 等ICLR 2024 · 被引用 27 次
- B2Opt: Learning to Optimize Black-box Optimization with Little BudgetXiaobin Li, Kai Wu, Xiaoyu Zhang, Handing WangAAAI 2025 · 被引用 23 次
它引用的顶会 Paper13
- Perceiver: General Perception with Iterative AttentionAndrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals 等ICML 2021 · 被引用 1,399 次
- Discovering Reinforcement Learning AlgorithmsJunhyuk Oh, Matteo Hessel, Wojciech M. Czarnecki, Zhongwen Xu 等NeurIPS 2020 · 被引用 154 次
- Discovered Policy OptimisationChris Lu, Jakub Grudzien Kuba, Alistair Letcher, Luke Metz 等NeurIPS 2022 · 被引用 134 次
- Improving Generalization in Meta Reinforcement Learning using Learned ObjectivesLouis Kirsch, Sjoerd van Steenkiste, Jürgen SchmidhuberICLR 2020 · 被引用 132 次
- A Self-Tuning Actor-Critic AlgorithmTom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel 等NeurIPS 2020 · 被引用 106 次
相关 Paper
- Discovering Temporally-Aware Reinforcement Learning AlgorithmsMatthew Thomas Jackson, Chris Lu, Louis Kirsch, Robert Tjarko Lange 等ICLR 2024 · 被引用 23 次
- Task-free Adaptive Meta Black-box OptimizationChao Wang, Licheng Jiao, Lingling Li, Jiaxuan Zhao 等ICLR 2026 · 被引用 4 次
- Neural Exploratory Landscape Analysis for Meta-Black-Box-OptimizationZeyuan Ma, Jiacheng Chen, Hongshu Guo, Yue-Jiao GongICLR 2025
- Low-Variance Gradient Estimation in Unrolled Computation Graphs with ES-SinglePaul VicolICML 2023 · 被引用 8 次
- Learning Discrete Structured Variational Auto-Encoder using Natural Evolution StrategiesAlon Berliner, Guy Rotman, Yossi Adi, Roi Reichart 等ICLR 2022 · 被引用 5 次
