Generalized Gradient Norm Clipping & Non-Euclidean -Smoothness
Thomas Pethick, Wanyun Xie, Mete Erdogan, Kimon Antonakopoulos, Antonio Silveti-Falls, Volkan Cevher
Abstract
This work introduces a hybrid non-Euclidean optimization method which generalizes gradient norm clipping by combining steepest descent and conditional gradient approaches. The method achieves the best of both worlds by establishing a descent property under a generalized notion of (L 0 ,L 1 )-smoothness. Weight decay is incorporated in a principled manner by identifying a connection to the Frank-Wolfe short step. In the stochastic case, we show an order optimal O(n -1/4 ) convergence rate by leveraging a momentum based gradient estimator. We discuss how to instantiate the algorithms for deep learning, which we dub Clipped Scion, and demonstrate their properties on image classification and language modeling. The code is available at https://github.com/LIONS-EPFL/ClippedScion. * Equal contribution. 2 By conditional gradient based methods, we mean those methods which leverage a linear minimization oracle lmo(d) = arg min x∈D ⟨d, x⟩ when updating their parameters with an open-loop stepsize. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Error Feedback for Muon and FriendsKaja Gruntkowska, Alexander Gaponov, Zhirayr Tovmasyan, Peter RichtárikICLR 2026 · 13 citations
- Nonlinearly Preconditioned Gradient Methods: Momentum and Stochastic AnalysisKonstantinos A. Oikonomidis, Jan Quan, Panagiotis PatrinosNeurIPS 2025 · 6 citations
- On the Interaction of Batch Noise, Adaptivity, and Compression, under -Smoothness: An SDE ApproachEnea Monzio Compagnoni, Rustem Islamov, Frank Proske, Aurelien Lucchi et al.ICML 2026 · 4 citations
- The Implicit Bias of Steepest Descent with Mini-batch Stochastic GradientJichu Li, Xuan Tang, Difan ZouICML 2026 · 1 citation
- Taming Stochastic Gradient Descent: Almost Sure Convergence and Saddle-Point Avoidance under -SmoothnessVassilis Apidopoulos, Iosif Lytras, Panayotis MertikopoulosICML 2026
Builds on17
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- Searching for Efficient Transformers for Language ModelingDavid R. So, Wojciech Manke, Hanxiao Liu, Zihang Dai et al.NeurIPS 2021 · 205 citations
- Momentum Improves Normalized SGDAshok Cutkosky, Harsh MehtaICML 2020 · 177 citations
- Improved Analysis of Clipping Algorithms for Non-convex OptimizationBohang Zhang, Jikai Jin, Cong Fang, Liwei WangNeurIPS 2020 · 139 citations
Related papers
- Stacey: Promoting Stochastic Steepest Descent via Accelerated ℓp-Smooth Nonconvex OptimizationXinyu Luo, Site Bai, Bolian Li, Petros Drineas et al.ICML 2025
- An Exploration of Non-Euclidean Gradient Descent: Muon and its Many VariantsMichael Crawshaw, Chirag Modi, Mingrui Liu, Robert GowerICML 2026 · 24 citations
- High-probability Bounds for Non-Convex Stochastic Optimization with Heavy TailsAshok Cutkosky, Harsh MehtaNeurIPS 2021 · 119 citations
- Convex and Non-convex Optimization Under Generalized SmoothnessHaochuan Li, Jian Qian, Yi Tian, Alexander Rakhlin et al.NeurIPS 2023 · 93 citations
- Methods for Convex (L0, L1)-Smooth Optimization: Clipping, Acceleration, and AdaptivityEduard Gorbunov, Nazarii Tupitsa, Sayantan Choudhury, Alen Aliev et al.ICLR 2025
