Quadratic models for understanding catapult dynamics of neural networks
Libin Zhu, Chaoyue Liu, Adityanarayanan Radhakrishnan, Mikhail Belkin
Abstract
While neural networks can be approximated by linear models as their width increases, certain properties of wide neural networks cannot be captured by linear models. In this work we show that recently proposed Neural Quadratic Models can exhibit the "catapult phase" (Lewkowycz et al., 2020) that arises when training such models with large learning rates. We then empirically show that the behaviour of neural quadratic models parallels that of neural networks in generalization, especially in the catapult phase regime. Our analysis further demonstrates that quadratic models can be an effective tool for analysis of neural networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68f20576-fe80-42da-9c2e-9bb7e74e4ac3Cited by top-tier papers4
- Catapults in SGD: spikes in the training loss and their impact on generalization through feature learningLibin Zhu, Chaoyue Liu, Adityanarayanan Radhakrishnan, Mikhail BelkinICML 2024 · 29 citations
- Closed-Form Training Dynamics Reveal Learned Features and Linear Structure in Word2Vec-like ModelsDhruva Karkada, James B. Simon, Yasaman Bahri, Michael R. DeWeeseNeurIPS 2025 · 9 citations
- Variational Learning Finds Flatter Solutions at the Edge of StabilityAvrajit Ghosh, Bai Cong, Rio Yokota, Saiprasad Ravishankar et al.NeurIPS 2025 · 2 citations
- Dropout Universality: Scaling Laws and Optimal Scheduling at the Edge-of-ChaosLucas Fernandez-SarmientoICML 2026
Builds on6
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent KernelStanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani et al.NeurIPS 2020 · 255 citations
- Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networksZiwei Ji, Matus TelgarskyICLR 2020 · 193 citations
- On the linearity of large non-linear models: when and why the tangent kernel is constantChaoyue Liu, Libin Zhu, Mikhail BelkinNeurIPS 2020 · 183 citations
- Dynamics of Deep Neural Networks and Neural Tangent HierarchyJiaoyang Huang, Horng-Tzer YauICML 2020 · 167 citations
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural NetworksYu Bai, Jason D. LeeICLR 2020 · 128 citations
Related papers
- Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning RegimeLeonardo Defilippis, Yizhou Xu, Julius Girardin, Vittorio Erba et al.ICLR 2026 · 20 citations
- Transition to Linearity of Wide Neural Networks is an Emerging Property of Assembling Weak ModelsChaoyue Liu, Libin Zhu, Mikhail BelkinICLR 2022 · 6 citations
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of GeneralizationBen Adlam, Jeffrey PenningtonICML 2020 · 133 citations
- The Onset of Variance-Limited Behavior for Networks in the Lazy and Rich RegimesAlexander B. Atanasov, Blake Bordelon, Sabarish Sainathan, Cengiz PehlevanICLR 2023 · 4 citations
- What can linearized neural networks actually say about generalization?Guillermo Ortiz-Jiménez, Seyed-Mohsen Moosavi-Dezfooli, Pascal FrossardNeurIPS 2021 · 62 citations
