Why Deep Jacobian Spectra Separate: Depth-Induced Scaling and Singular-Vector Alignment
Nathanaël Haas, François Gatine, Augustin Cosse, Zied Bouraoui
Abstract
Understanding why gradient-based training in deep networks exhibits strong implicit bias remains challenging, in part because tractable singular-value dynamics are typically available only for balanced deep linear models. We propose an alternative route based on two theoretically grounded and empirically testable signatures of deep Jacobians: depth-induced exponential scaling of ordered singular values and strong spectral separation. Adopting a fixed-gates view of piecewise-linear networks, where Jacobians reduce to products of masked linear maps within a single activation region, we prove the existence of Lyapunov exponents governing the top singular values at initialization, give closed-form expressions in a tractable masked model, and quantify finite-depth corrections. We further show that sufficiently strong separation forces singular-vector alignment in matrix products, yielding an approximately shared singular basis for intermediate Jacobians. Together, these results motivate an approximation regime in which singular-value dynamics become effectively decoupled, mirroring classical balanced deep-linear analyses without requiring balancing. Experiments in fixed-gates settings validate the predicted scaling, alignment, and resulting dynamics, supporting a mechanistic account of emergent low-rank Jacobian structure as a driver of implicit bias.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 226 citations
- Nuclear Norm Regularization for Deep LearningChristopher Scarvelis, Justin M. SolomonNeurIPS 2024 · 17 citations
- Generalization or Hallucination? Understanding Out-of-Context Reasoning in TransformersYixiao Huang, Hanlin Zhu, Tianyu Guo, Jiantao Jiao et al.NeurIPS 2025 · 10 citations
- Implicit Bias of Large Depth Networks: a Notion of Rank for Nonlinear FunctionsArthur JacotICLR 2023 · 2 citations
Related papers
- Training invariances and the low-rank phenomenon: beyond linear networksThien Le, Stefanie JegelkaICLR 2022 · 39 citations
- The Implicit Bias of Depth: From Neural Collapse to Softmax CodesConnall Garrod, Jonathan Keating, Christos ThrampoulidisICML 2026
- Saddle-To-Saddle Dynamics in Deep ReLU Networks: Low-Rank Bias in the First Saddle EscapeIoannis Bantzis, James B. Simon, Arthur JacotICLR 2026 · 4 citations
- On the spectral bias of two-layer linear networksAditya Vardhan Varre, Maria-Luiza Vladarean, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2023 · 27 citations
- Unique Properties of Flat Minima in Deep NetworksRotem Mulayoff, Tomer MichaeliICML 2020 · 43 citations
