Understanding the Mechanisms of Fast Hyperparameter Transfer
Nikhil Ghosh, Denny Wu, Alberto Bietti
Abstract
The growing scale of deep learning models has rendered standard hyperparameter (HP) optimization prohibitively expensive. A promising solution is the use of scale-aware hyperparameters, which can enable direct transfer of optimal HPs from small-scale grid searches to large models with minimal performance loss. To understand the principles governing such transfer strategy, we develop a general conceptual framework for reasoning about HP transfer across scale, characterizing transfer as fast when the suboptimality it induces vanishes asymptotically faster than the finite-scale performance gap. We show formally that fast transfer is equivalent to useful transfer for compute-optimal grid search, meaning that transfer is asymptotically more compute-efficient than direct tuning. While empirical work has found that the Maximal Update Parameterization (P) exhibits fast transfer when scaling model width, the mechanisms remain poorly understood. We show that this property depends critically on problem structure by presenting synthetic settings where transfer either offers provable computational advantage or fails to outperform direct tuning even under P. To explain the fast transfer observed in practice, we conjecture that decomposing the optimization trajectory reveals two contributions to loss reduction: (1) a width-stable component that determines the optimal HPs, and (2) a width-sensitive component that improves with width but weakly perturbs the HP optimum. We present empirical evidence for this hypothesis across various settings, including large language model pretraining.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b8bf6e77-65ea-4242-9276-01713defd331Cited by top-tier papers1
Ask how each one uses itBuilds on15
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 242 citations
- Generalisation error in learning with random features and the hidden manifold modelFederica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard et al.ICML 2020 · 184 citations
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang et al.NeurIPS 2022 · 173 citations
- Tensor Programs VI: Feature Learning in Infinite Depth Neural NetworksGreg Yang, Dingli Yu, Chen Zhu, Soufiane HayouICLR 2024 · 77 citations
- Depthwise Hyperparameter Transfer in Residual Networks: Dynamics and Scaling LimitBlake Bordelon, Lorenzo Noci, Mufan Bill Li, Boris Hanin et al.ICLR 2024 · 54 citations
Related papers
- On the Provable Separation of Scales in Maximal Update ParameterizationLetong Hong, Zhangyang WangICML 2025
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor et al.NeurIPS 2021 · 208 citations
- Hyperparameter Transfer Laws for Non-Recurrent Multi-Path Neural NetworksHaosong Zhang, Shenxi Wu, Xingjian Ma, Shirui Bian et al.ICML 2026 · 1 citation
- Weight Decay may matter more than µP for Learning Rate Transfer in PracticeAtli Kosson, Jeremy Welborn, Yang Liu, Martin Jaggi et al.ICLR 2026 · 11 citations
- Completed Hyperparameter Transfer across Modules, Width, Depth, Batch and DurationBruno Mlodozeniec, Pierre Ablin, Louis Béthune, Dan Busbridge et al.ICLR 2026 · 24 citations
