Learning by Directional Gradient Descent
David Silver, Anirudh Goyal, Ivo Danihelka, Matteo Hessel, Hado van Hasselt
摘要
How should state be constructed from a sequence of observations, so as to best achieve some objective? Most deep learning methods update the parameters of the state representation by gradient descent. However, no prior method for computing the gradient is fully satisfactory, for example consuming too much memory, introducing too much variance, or adding too much bias. In this work, we propose a new learning algorithm that addresses these limitations. The basic idea is to update the parameters of the representation by using the directional derivative along a candidate direction, a quantity that may be computed online with the same computational cost as the representation itself. We consider several different choices of candidate direction, including random selection and approximations to the true gradient, and investigate their performance on several synthetic tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Revisiting Zeroth-Order Optimization for Memory-Efficient LLM Fine-Tuning: A BenchmarkYihua Zhang, Pingzhi Li, Junyuan Hong, Jiaxiang Li 等ICML 2024 · 被引用 134 次
- DeepZero: Scaling Up Zeroth-Order Optimization for Deep Model TrainingAochuan Chen, Yimeng Zhang, Jinghan Jia, James Diffenderfer 等ICLR 2024 · 被引用 88 次
- Online learning of long-range dependenciesNicolas Zucchet, Robert Meier, Simon Schug, Asier Mujika 等NeurIPS 2023 · 被引用 43 次
- Can Forward Gradient Match Backpropagation?Louis Fournier, Stéphane Rivaud, Eugene Belilovsky, Michael Eickenberg 等ICML 2023 · 被引用 33 次
- DPZero: Private Fine-Tuning of Language Models without BackpropagationLiang Zhang, Bingcong Li, Kiran Koshy Thekumparampil, Sewoong Oh 等ICML 2024 · 被引用 27 次
它引用的顶会 Paper2
- Unbiased Gradient Estimation in Unrolled Computation Graphs with Persistent Evolution StrategiesPaul Vicol, Luke Metz, Jascha Sohl-DicksteinICML 2021 · 被引用 77 次
- Learned Initializations for Optimizing Coordinate-Based Neural RepresentationsMatthew Tancik, Ben Mildenhall, Terrance Wang, Divi Schmidt 等CVPR 2021
相关 Paper
- Meta-Gradient Reinforcement Learning with an Objective Discovered OnlineZhongwen Xu, Hado Philip van Hasselt, Matteo Hessel, Junhyuk Oh 等NeurIPS 2020 · 被引用 90 次
- Bootstrapped Representations in Reinforcement LearningCharline Le Lan, Stephen Tu, Mark Rowland, Anna Harutyunyan 等ICML 2023 · 被引用 12 次
- Unbiased learning of deep generative models with structured discrete representationsHenry C. Bendekgey, Gabe Hope, Erik B. SudderthNeurIPS 2023 · 被引用 2 次
- Training Recurrent Neural Networks Online by Learning Explicit State VariablesSomjit Nath, Vincent Liu, Alan Chan, Xin Li 等ICLR 2020 · 被引用 9 次
- Learning Belief Representations for Partially Observable Deep RLAndrew Wang, Andrew C. Li, Toryn Q. Klassen, Rodrigo Toro Icarte 等ICML 2023 · 被引用 21 次
