How many degrees of freedom do we need to train deep networks: a loss landscape perspective
Brett W. Larsen, Stanislav Fort, Nic Becker, Surya Ganguli
摘要
A variety of recent works, spanning pruning, lottery tickets, and training within random subspaces, have shown that deep neural networks can be trained using far fewer degrees of freedom than the total number of parameters. We analyze this phenomenon for random subspaces by first examining the success probability of hitting a training loss sub-level set when training within a random subspace of a given training dimensionality. We find a sharp phase transition in the success probability from to as the training dimension surpasses a threshold. This threshold training dimension increases as the desired final loss decreases, but decreases as the initial loss decreases. We then theoretically explain the origin of this phase transition, and its dependence on initialization and final desired loss, in terms of properties of the high-dimensional geometry of the loss landscape. In particular, we show via Gordon's escape theorem, that the training dimension plus the Gaussian width of the desired loss sub-level set, projected onto a unit sphere surrounding the initialization, must exceed the total number of parameters for the success probability to be large. In several architectures and datasets, we measure the threshold training dimension as a function of initialization and demonstrate that it is a small fraction of the total parameters, implying by our theory that successful training with so few dimensions is possible precisely because the Gaussian width of low loss sub-level sets is very large. Moreover, we compare this threshold training dimension to more sophisticated ways of reducing training degrees of freedom, including lottery tickets as well as a new, analogous method: lottery subspaces. Code is available at https://github.com/ganguli-lab/degrees-of-freedom.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang 等ICML 2024 · 被引用 433 次
- Memory-Efficient LLM Training with Online Subspace DescentKaizhao Liang, Bo Liu, Lizhang Chen, Qiang LiuNeurIPS 2024 · 被引用 46 次
- VeLoRA: Memory Efficient Training using Rank-1 Sub-Token ProjectionsRoy Miles, Pradyumna Reddy, Ismail Elezi, Jiankang DengNeurIPS 2024 · 被引用 22 次
- LoQT: Low-Rank Adapters for Quantized PretrainingSebastian Loeschcke, Mads Toftrup, Michael J. Kastoryano, Serge J. Belongie 等NeurIPS 2024 · 被引用 14 次
- How Sparse Can We Prune A Deep Network: A Fundamental Limit PerspectiveQiaozhe Zhang, Ruijie Zhang, Jun Sun, Yingzhuang LiuNeurIPS 2024 · 被引用 14 次
它引用的顶会 Paper4
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 被引用 884 次
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 被引用 743 次
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent KernelStanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani 等NeurIPS 2020 · 被引用 255 次
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit 等ICLR 2020 · 被引用 198 次
相关 Paper
- Polynomially Over-Parameterized Convolutional Neural Networks Contain Structured Strong Winning Lottery TicketsArthur da Cunha, Francesco d'Amore, Emanuele NataleNeurIPS 2023 · 被引用 5 次
- On the Existence of Universal Lottery TicketsRebekka Burkholz, Nilanjana Laha, Rajarshi Mukherjee, Alkis GotovosICLR 2022 · 被引用 38 次
- Dual Lottery Ticket HypothesisYue Bai, Huan Wang, Zhiqiang Tao, Kunpeng Li 等ICLR 2022 · 被引用 49 次
- Logarithmic Pruning is All You NeedLaurent Orseau, Marcus Hutter, Omar RivasplataNeurIPS 2020 · 被引用 102 次
- Validating the Lottery Ticket Hypothesis with Inertial Manifold TheoryZeru Zhang, Jiayin Jin, Zijie Zhang, Yang Zhou 等NeurIPS 2021 · 被引用 45 次
