Deep Hierarchical Learning with Nested Subspace Networks for Large Language Models
Paulius Rauba, Mihaela van der Schaar
摘要
Large neural networks are typically trained for a fixed computational budget, creating a rigid trade-off between performance and efficiency that is ill-suited for deployment in resource-constrained or dynamic environments. Existing approaches to this problem present a difficult choice: training a discrete collection of specialist models is computationally prohibitive, while dynamic methods like slimmable networks often lack the flexibility to be applied to large, pre-trained foundation models. In this work, we propose Nested Subspace Networks (NSNs), a novel architectural paradigm that enables a single model to be dynamically and granularly adjusted across a continuous spectrum of compute budgets at inference time. The core of our approach is to re-parameterize linear layers to satisfy a nested subspace property, such that the function computed at a given rank is a strict subspace of the function at any higher rank. We show that this entire hierarchy of models can be optimized jointly via an uncertainty-aware objective that learns to balance the contributions of different ranks based on their intrinsic difficulty. We demonstrate empirically that NSNs can be surgically applied to pre-trained LLMs and unlock a smooth and predictable compute-performance frontier. For example, a single NSN-adapted model can achieve a 50% reduction in inference FLOPs with only a 5 percentage point loss in accuracy. Our findings establish NSNs as a powerful framework for creating the next generation of adaptive foundation models. On the other hand, recent methods using dynamic neural networks (Han et al., 2021) operate by designing architectures that can be adjusted at inference time, such as slimmable networks that can drop channels (Yu et al., 2018; Li et al., 2021) or layers (Wu et al., 2018) . In theory, these approaches more readily take advantage of a single set of weights to serve multiple budgets. In practice, however, this strategy often comes at the price of much more challenging, specialized training schemes that
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper18
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 被引用 695 次
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 被引用 656 次
- Universally Slimmable Networks and Improved Training TechniquesJiahui Yu, Thomas S. HuangICCV 2019 · 被引用 444 次
相关 Paper
- Bayesian Nested Neural Networks for Uncertainty Calibration and Adaptive CompressionYufei Cui, Ziquan Liu, Qiao Li, Antoni B. Chan 等CVPR 2021
- Balcony: A Lightweight Approach to Dynamic Inference of Generative Language ModelsBenyamin Jamialahmadi, Parsa Kavehzadeh, Mehdi Rezagholizadeh, Parsa Farinneya 等EMNLP 2025
- Star Elastic: Many-in-One Reasoning LLMs with Efficient Budget ControlAli Taghibakhshi, Ruisi Cai, Saurav Muralidharan, Sharath Turuvekere Sreenivas 等ICML 2026 · 被引用 1 次
- Navigating Scaling Laws: Compute Optimality in Adaptive Model TrainingSotiris Anagnostidis, Gregor Bachmann, Imanol Schlag, Thomas HofmannICML 2024 · 被引用 2 次
- NestedFP: High-Performance, Memory-Efficient Dual-Precision Floating Point Support for LLMsHaeun Lee, Omin Kwon, Yeonhong Park, Jae W. LeeNeurIPS 2025 · 被引用 5 次
