Revisiting Scalable Hessian Diagonal Approximations for Applications in Reinforcement Learning
Mohamed Elsayed, Homayoon Farrahi, Felix Dangel, A. Rupam Mahmood
摘要
Second-order information is valuable for many applications but challenging to compute. Several works focus on computing or approximating Hessian diagonals, but even this simplification introduces significant additional costs compared to computing a gradient. In the absence of efficient exact computation schemes for Hessian diagonals, we revisit an early approximation scheme proposed by Becker and LeCun (1989, BL89), which has a cost similar to gradients and appears to have been overlooked by the community. We introduce HesScale, an improvement over BL89, which adds negligible extra computation. On small networks, we find that this improvement is of higher quality than all alternatives, even those with theoretical guarantees, such as unbiasedness, while being much cheaper to compute. We use this insight in reinforcement learning problems where small networks are used and demonstrate HesScale in second-order optimization and scaling the step-size parameter. In our experiments, HesScale optimizes faster than existing methods and improves stability through step-size scaling. These findings are promising for scaling secondorder methods in larger models in the future. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Convolutions and More as Einsum: A Tensor Network Perspective with Advances for Second-Order MethodsFelix DangelNeurIPS 2024 · 被引用 5 次
- Mitigating Plasticity Loss in Continual Reinforcement Learning by Reducing ChurnHongyao Tang, Johan S. Obando-Ceron, Pablo Samuel Castro, Aaron C. Courville 等ICML 2025
- Reinforcement Learning with Adaptive Reward Modeling for Expensive-to-Evaluate SystemsHongyuan Su, Yu Zheng, Yuan Yuan, Yuming Lin 等ICML 2025
- AdaMeZO: Adam-style Zeroth-Order Optimizer for LLM Fine-tuning Without Maintaining the MomentsZhijie Cai, Haolong Chen, Guangxu ZhuICML 2026
它引用的顶会 Paper7
- ADAHESSIAN: An Adaptive Second Order Optimizer for Machine LearningZhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa 等AAAI 2021 · 被引用 358 次
- WoodFisher: Efficient Second-Order Approximation for Neural Network CompressionSidak Pal Singh, Dan AlistarhNeurIPS 2020 · 被引用 217 次
- BackPACK: Packing more into BackpropFelix Dangel, Frederik Kunstner, Philipp HennigICLR 2020 · 被引用 114 次
- Variational Learning is Effective for Large Deep NetworksYuesong Shen, Nico Daheim, Bai Cong, Peter Nickl 等ICML 2024 · 被引用 53 次
- Addressing Loss of Plasticity and Catastrophic Forgetting in Continual LearningMohamed Elsayed, A. Rupam MahmoodICLR 2024 · 被引用 52 次
相关 Paper
- SOSP: Efficiently Capturing Global Correlations by Second-Order Structured PruningManuel Nonnenmacher, Thomas Pfeil, Ingo Steinwart, David ReebICLR 2022 · 被引用 48 次
- Sassha: Sharpness-aware Adaptive Second-order Optimization with Stable Hessian ApproximationDahun Shin, Dongyeop Lee, Jinseok Chung, Namhoon LeeICML 2025
- SPAN: A Stochastic Projected Approximate Newton MethodXunpeng Huang, Xianfeng Liang, Zhengyang Liu, Lei Li 等AAAI 2020 · 被引用 4 次
- ISAAC Newton: Input-based Approximate Curvature for Newton's MethodFelix Petersen, Tobias Sutter, Christian Borgelt, Dongsung Huh 等ICLR 2023
- Unifying Gradient Estimators for Meta-Reinforcement Learning via Off-Policy EvaluationYunhao Tang, Tadashi Kozuno, Mark Rowland, Rémi Munos 等NeurIPS 2021 · 被引用 9 次
