Gradient Smoothing: Coupling Layer-wise Updates for Improved Optimization
Haoming Meng, Anton Sugolov, Vardan Papyan
Abstract
Deep neural networks with repeated architectural blocks, such as transformers, often exhibit structured relationships across layers that emerge during training. Motivated by this observation, we introduce Depth-wise Gradient Augmentation, a general optimization paradigm in which the update applied to each layer is obtained by transforming the collection of block-wise optimizer updates along the depth dimension. Within this framework, we study Gradient Smoothing, a family of depth-wise smoothing methods, and instantiate it with a simple local Window Smoothing operator. The resulting method operates directly on block-wise updates produced by arbitrary base optimizers (e.g., SGD, Adam, Muon), incurs minimal computational overhead, and is compatible with existing optimization pipelines. We evaluate Gradient Smoothing across a diverse set of architectures and training regimes, including language model pretraining, RL post-training of LLMs for reasoning, diffusion modeling, and image classification with Vision Transformers. Across these settings, Gradient Smoothing consistently improves optimization and generalization performance without modifying model architectures or training objectives. We further show that it promotes more structured representation evolution across depth, consistent with its interpretation as a structured depth-wise preconditioning method. Together, these results establish Depth-wise Gradient Augmentation as a promising framework for exploiting cross-depth structure in optimization and demonstrate Gradient Smoothing as a simple and broadly applicable instantiation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1fd27b97-fa36-473e-95fc-3bbb31c3a732Builds on19
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Transformers Represent Belief State Geometry in their Residual StreamAdam S. Shai, Lucas Teixeira, Alexander Gietelink Oldenziel, Sarah Marzen et al.NeurIPS 2024 · 83 citations
- Kronecker-Factored Approximate Curvature for Modern Neural Network ArchitecturesRuna Eschenhagen, Alexander Immer, Richard E. Turner, Frank Schneider et al.NeurIPS 2023 · 62 citations
- Deep Neural Collapse Is Provably Optimal for the Deep Unconstrained Features ModelPeter Súkeník, Marco Mondelli, Christoph H. LampertNeurIPS 2023 · 51 citations
- Linguistic Collapse: Neural Collapse in (Large) Language ModelsRobert Wu, Vardan PapyanNeurIPS 2024 · 45 citations
Related papers
- Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across ScalesShikai Qiu, Charlie Chen, Hoang Phan, Qi Lei et al.NeurIPS 2025 · 17 citations
- Efficient Parallel Samplers for Recurrent-Depth ModelsJonas Geiping, Xinyu Yang, Guinan SuICML 2026 · 5 citations
- Mitigating Over-smoothing in Transformers via Regularized Nonlocal FunctionalsTam Nguyen, Tan M. Nguyen, Richard G. BaraniukNeurIPS 2023 · 46 citations
- Extrapolation for Large-batch Training in Deep LearningTao Lin, Lingjing Kong, Sebastian U. Stich, Martin JaggiICML 2020 · 43 citations
- AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-trainingHuishuai Zhang, Bohan Wang, Luoxin ChenEMNLP 2025 · 1 citation
