BSFA: Leveraging the Subspace Dichotomy to Accelerate Neural Network Training
Wenjie Zhou, Bohan Wang, Wei Chen, Xueqi Cheng
Abstract
Recent studies (Gur-Ari et al., 2018; Song et al., 2024; Wen et al., 2024) highlight a fundamental dichotomy in deep learning optimization: Although parameter updates along the top eigendirections of the loss Hessian (Domspace) capture most of the update magnitude, they often contribute minimally to loss reduction. In contrast, updates in the orthogonal component (Bulk-space) have smaller magnitudes but drive most learning progress. In this work, we further advance the understanding of this phenomenon and introduce the Bulk-Space-Filtration-Accelerator (BSFA), a novel plugand-play framework. BSFA accelerates training by differentially scaling update components projected onto these distinct subspaces, simultaneously enhancing stability by moderating updates in the dominant subspace and boosting convergence speed by amplifying those in the bulk-space. To ensure BSFA is both practical and scalable for contemporary large models, we introduce two key innovations: an efficient estimator using Principal Component Analysis (PCA) on historical updates for fast subspace estimation, and a block-wise strategy that applies this estimation on a per-parameter-block basis. These designs make BSFA computationally tractable and highly effective. We demonstrate BSFA's acceleration across various tasks, notably achieving approximately 2× speedup when pre-training LLaMA-72M on WikiText-103 and LLaMA-134M on OpenWebText compared to vanilla AdamW.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real et al.NeurIPS 2023 · 734 citations
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-trainingHong Liu, Zhiyuan Li, David Leo Wright Hall, Percy Liang et al.ICLR 2024 · 264 citations
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit et al.ICLR 2020 · 198 citations
Related papers
- Does SGD really happen in tiny subspaces?Minhak Song, Kwangjun Ahn, Chulhee YunICLR 2025
- PRAC: Principal-Random Subspace for LLM Activation Compression and Memory-Efficient TrainingYanyi Li, Yimu Zhang, Cong FangICML 2026
- Memory-Efficient LLM Training with Online Subspace DescentKaizhao Liang, Bo Liu, Lizhang Chen, Qiang LiuNeurIPS 2024 · 46 citations
- The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-TrainingJinbo Wang, Mingze Wang, Zhanpeng Zhou, Junchi Yan et al.ICML 2025
- SUMO: Subspace-Aware Moment-Orthogonalization for Accelerating Memory-Efficient LLM TrainingYehonathan Refael, Guy Smorodinsky, Tom Tirer, Ofir LindenbaumNeurIPS 2025 · 17 citations
