Low-Variance Gradient Estimation in Unrolled Computation Graphs with ES-Single
Paul Vicol
Abstract
We propose an evolution strategies-based algorithm for estimating gradients in unrolled computation graphs, called ES-Single. Similarly to the recently-proposed Persistent Evolution Strategies (PES), ES-Single is unbiased, and overcomes chaos arising from recursive function applications by smoothing the meta-loss landscape. ES-Single samples a single perturbation per particle, that is kept fixed over the course of an inner problem (e.g., perturbations are not re-sampled for each partial unroll). Compared to PES, ES-Single is simpler to implement and has lower variance: the variance of ES-Single is constant with respect to the number of truncated unrolls, removing a key barrier in applying ES to long inner problems using short truncations. We show that ES-Single is unbiased for quadratic inner problems, and demonstrate empirically that its variance can be substantially lower than that of PES. ES-Single consistently outperforms PES on a variety of tasks, including a synthetic benchmark task, hyperparameter optimization, training recurrent neural networks, and training learned optimizers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a8612bf4-bfea-4a70-8de5-48bd91f203ddCited by top-tier papers5
- Revisiting Zeroth-Order Optimization for Memory-Efficient LLM Fine-Tuning: A BenchmarkYihua Zhang, Pingzhi Li, Junyuan Hong, Jiaxiang Li et al.ICML 2024 · 134 citations
- DeepZero: Scaling Up Zeroth-Order Optimization for Deep Model TrainingAochuan Chen, Yimeng Zhang, Jinghan Jia, James Diffenderfer et al.ICLR 2024 · 88 citations
- Variance-Reduced Gradient Estimation via Noise-Reuse in Online Evolution StrategiesOscar Li, James Harrison, Jascha Sohl-Dickstein, Virginia Smith et al.NeurIPS 2023 · 11 citations
- μLO: Compute-Efficient Meta-Generalization of Learned OptimizersBenjamin Thérien, Charles-Étienne Joseph, Boris Knyazev, Edouard Oyallon et al.ICLR 2026 · 10 citations
- Forward-Only Continual LearningJiao Chen, Jiayi He, Fangfang Chen, Zuohong Lv et al.ACM MM 2025 · 1 citation
Builds on6
- Efficient and Modular Implicit DifferentiationMathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig et al.NeurIPS 2022 · 386 citations
- Unbiased Gradient Estimation in Unrolled Computation Graphs with Persistent Evolution StrategiesPaul Vicol, Luke Metz, Jascha Sohl-DicksteinICML 2021 · 77 citations
- Gradient Descent: The Ultimate OptimizerKartik Chandra, Audrey Xie, Jonathan Ragan-Kelley, Erik MeijerNeurIPS 2022 · 66 citations
- On Implicit Bias in Overparameterized Bilevel OptimizationPaul Vicol, Jonathan P. Lorraine, Fabian Pedregosa, David Duvenaud et al.ICML 2022 · 48 citations
- Learning by Directional Gradient DescentDavid Silver, Anirudh Goyal, Ivo Danihelka, Matteo Hessel et al.ICLR 2022 · 44 citations
Related papers
- ES-MAML: Simple Hessian-Free Meta LearningXingyou Song, Wenbo Gao, Yuxiang Yang, Krzysztof Choromanski et al.ICLR 2020 · 128 citations
- Discovering Evolution Strategies via Meta-Black-Box OptimizationRobert Tjarko Lange, Tom Schaul, Yutian Chen, Tom Zahavy et al.ICLR 2023 · 21 citations
- Evolution Strategies at the HyperscaleBidipta Sarkar, Mattie Fellows, Juan Duque, Alistair Letcher et al.ICML 2026 · 16 citations
- Learning Discrete Structured Variational Auto-Encoder using Natural Evolution StrategiesAlon Berliner, Guy Rotman, Yossi Adi, Roi Reichart et al.ICLR 2022 · 5 citations
- Generalizing Gaussian Smoothing for Random SearchKatelyn Gao, Ozan SenerICML 2022 · 22 citations
