Revisiting Scalable Hessian Diagonal Approximations for Applications in Reinforcement Learning
Mohamed Elsayed, Homayoon Farrahi, Felix Dangel, A. Rupam Mahmood
Abstract
Second-order information is valuable for many applications but challenging to compute. Several works focus on computing or approximating Hessian diagonals, but even this simplification introduces significant additional costs compared to computing a gradient. In the absence of efficient exact computation schemes for Hessian diagonals, we revisit an early approximation scheme proposed by Becker and LeCun (1989, BL89), which has a cost similar to gradients and appears to have been overlooked by the community. We introduce HesScale, an improvement over BL89, which adds negligible extra computation. On small networks, we find that this improvement is of higher quality than all alternatives, even those with theoretical guarantees, such as unbiasedness, while being much cheaper to compute. We use this insight in reinforcement learning problems where small networks are used and demonstrate HesScale in second-order optimization and scaling the step-size parameter. In our experiments, HesScale optimizes faster than existing methods and improves stability through step-size scaling. These findings are promising for scaling secondorder methods in larger models in the future. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4acce4d2-b5f6-4ddf-91b5-586ddbeefeeaCited by top-tier papers4
- Convolutions and More as Einsum: A Tensor Network Perspective with Advances for Second-Order MethodsFelix DangelNeurIPS 2024 · 5 citations
- Mitigating Plasticity Loss in Continual Reinforcement Learning by Reducing ChurnHongyao Tang, Johan S. Obando-Ceron, Pablo Samuel Castro, Aaron C. Courville et al.ICML 2025
- Reinforcement Learning with Adaptive Reward Modeling for Expensive-to-Evaluate SystemsHongyuan Su, Yu Zheng, Yuan Yuan, Yuming Lin et al.ICML 2025
- AdaMeZO: Adam-style Zeroth-Order Optimizer for LLM Fine-tuning Without Maintaining the MomentsZhijie Cai, Haolong Chen, Guangxu ZhuICML 2026
Builds on7
- ADAHESSIAN: An Adaptive Second Order Optimizer for Machine LearningZhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa et al.AAAI 2021 · 358 citations
- WoodFisher: Efficient Second-Order Approximation for Neural Network CompressionSidak Pal Singh, Dan AlistarhNeurIPS 2020 · 217 citations
- BackPACK: Packing more into BackpropFelix Dangel, Frederik Kunstner, Philipp HennigICLR 2020 · 114 citations
- Variational Learning is Effective for Large Deep NetworksYuesong Shen, Nico Daheim, Bai Cong, Peter Nickl et al.ICML 2024 · 53 citations
- Addressing Loss of Plasticity and Catastrophic Forgetting in Continual LearningMohamed Elsayed, A. Rupam MahmoodICLR 2024 · 52 citations
Related papers
- SOSP: Efficiently Capturing Global Correlations by Second-Order Structured PruningManuel Nonnenmacher, Thomas Pfeil, Ingo Steinwart, David ReebICLR 2022 · 48 citations
- Sassha: Sharpness-aware Adaptive Second-order Optimization with Stable Hessian ApproximationDahun Shin, Dongyeop Lee, Jinseok Chung, Namhoon LeeICML 2025
- SPAN: A Stochastic Projected Approximate Newton MethodXunpeng Huang, Xianfeng Liang, Zhengyang Liu, Lei Li et al.AAAI 2020 · 4 citations
- ISAAC Newton: Input-based Approximate Curvature for Newton's MethodFelix Petersen, Tobias Sutter, Christian Borgelt, Dongsung Huh et al.ICLR 2023
- Unifying Gradient Estimators for Meta-Reinforcement Learning via Off-Policy EvaluationYunhao Tang, Tadashi Kozuno, Mark Rowland, Rémi Munos et al.NeurIPS 2021 · 9 citations
