Low-Variance Gradient Estimation in Unrolled Computation Graphs with ES-Single
Paul Vicol
摘要
We propose an evolution strategies-based algorithm for estimating gradients in unrolled computation graphs, called ES-Single. Similarly to the recently-proposed Persistent Evolution Strategies (PES), ES-Single is unbiased, and overcomes chaos arising from recursive function applications by smoothing the meta-loss landscape. ES-Single samples a single perturbation per particle, that is kept fixed over the course of an inner problem (e.g., perturbations are not re-sampled for each partial unroll). Compared to PES, ES-Single is simpler to implement and has lower variance: the variance of ES-Single is constant with respect to the number of truncated unrolls, removing a key barrier in applying ES to long inner problems using short truncations. We show that ES-Single is unbiased for quadratic inner problems, and demonstrate empirically that its variance can be substantially lower than that of PES. ES-Single consistently outperforms PES on a variety of tasks, including a synthetic benchmark task, hyperparameter optimization, training recurrent neural networks, and training learned optimizers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Revisiting Zeroth-Order Optimization for Memory-Efficient LLM Fine-Tuning: A BenchmarkYihua Zhang, Pingzhi Li, Junyuan Hong, Jiaxiang Li 等ICML 2024 · 被引用 134 次
- DeepZero: Scaling Up Zeroth-Order Optimization for Deep Model TrainingAochuan Chen, Yimeng Zhang, Jinghan Jia, James Diffenderfer 等ICLR 2024 · 被引用 88 次
- Variance-Reduced Gradient Estimation via Noise-Reuse in Online Evolution StrategiesOscar Li, James Harrison, Jascha Sohl-Dickstein, Virginia Smith 等NeurIPS 2023 · 被引用 11 次
- μLO: Compute-Efficient Meta-Generalization of Learned OptimizersBenjamin Thérien, Charles-Étienne Joseph, Boris Knyazev, Edouard Oyallon 等ICLR 2026 · 被引用 10 次
- Forward-Only Continual LearningJiao Chen, Jiayi He, Fangfang Chen, Zuohong Lv 等ACM MM 2025 · 被引用 1 次
它引用的顶会 Paper6
- Efficient and Modular Implicit DifferentiationMathieu Blondel, Quentin Berthet, Marco Cuturi, Roy Frostig 等NeurIPS 2022 · 被引用 386 次
- Unbiased Gradient Estimation in Unrolled Computation Graphs with Persistent Evolution StrategiesPaul Vicol, Luke Metz, Jascha Sohl-DicksteinICML 2021 · 被引用 77 次
- Gradient Descent: The Ultimate OptimizerKartik Chandra, Audrey Xie, Jonathan Ragan-Kelley, Erik MeijerNeurIPS 2022 · 被引用 66 次
- On Implicit Bias in Overparameterized Bilevel OptimizationPaul Vicol, Jonathan P. Lorraine, Fabian Pedregosa, David Duvenaud 等ICML 2022 · 被引用 48 次
- Learning by Directional Gradient DescentDavid Silver, Anirudh Goyal, Ivo Danihelka, Matteo Hessel 等ICLR 2022 · 被引用 44 次
相关 Paper
- ES-MAML: Simple Hessian-Free Meta LearningXingyou Song, Wenbo Gao, Yuxiang Yang, Krzysztof Choromanski 等ICLR 2020 · 被引用 128 次
- Discovering Evolution Strategies via Meta-Black-Box OptimizationRobert Tjarko Lange, Tom Schaul, Yutian Chen, Tom Zahavy 等ICLR 2023 · 被引用 21 次
- Evolution Strategies at the HyperscaleBidipta Sarkar, Mattie Fellows, Juan Duque, Alistair Letcher 等ICML 2026 · 被引用 16 次
- Learning Discrete Structured Variational Auto-Encoder using Natural Evolution StrategiesAlon Berliner, Guy Rotman, Yossi Adi, Roi Reichart 等ICLR 2022 · 被引用 5 次
- Generalizing Gaussian Smoothing for Random SearchKatelyn Gao, Ozan SenerICML 2022 · 被引用 22 次
