Improving Neural Network Training in Low Dimensional Random Bases
Frithjof Gressmann, Zach Eaton-Rosen, Carlo Luschi
Abstract
Stochastic Gradient Descent (SGD) has proven to be remarkably effective in optimizing deep neural networks that employ ever-larger numbers of parameters. Yet, improving the efficiency of large-scale optimization remains a vital and highly active area of research. Recent work has shown that deep neural networks can be optimized in randomly-projected subspaces of much smaller dimensionality than their native parameter space. While such training is promising for more efficient and scalable optimization schemes, its practical application is limited by inferior optimization performance. Here, we improve on recent random subspace approaches as follows: Firstly, we show that keeping the random projection fixed throughout training is detrimental to optimization. We propose re-drawing the random subspace at each step, which yields significantly better performance. We realize further improvements by applying independent projections to different parts of the network, making the approximation more efficient as network dimensionality grows. To implement these experiments, we leverage hardware-accelerated pseudo-random number generation to construct the random projections on-demand at every optimization step, allowing us to distribute the computation of independent random directions across multiple workers with shared random seeds. This yields significant reductions in memory and is up to 10 times faster for the workloads in question.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a1df38d7-ab1c-4911-9584-3aefa4eba6feCited by top-tier papers14
- Subspace Adversarial TrainingTao Li, Yingwen Wu, Sizhe Chen, Kun Fang et al.CVPR 2022 · 59 citations
- Subspace Learning for Effective Meta-LearningWeisen Jiang, James T. Kwok, Yu ZhangICML 2022 · 28 citations
- Fine-tuning Happens in Tiny Subspaces: Exploring Intrinsic Task-specific Subspaces of Pre-trained Language ModelsZhong Zhang, Bang Liu, Junming ShaoACL 2023 · 8 citations
- Lifelong Test-Time Adaptation via Online Learning in Tracked Low-Dimensional SubspaceDexin Duan, Rui Xu, Peilin Liu, Fei WenNeurIPS 2025 · 7 citations
- Identifying Policy Gradient SubspacesJan Schneider, Pierre Schumacher, Simon Guist, Le Chen et al.ICLR 2024 · 7 citations
Builds on1
Related papers
- Trainable Weight Averaging: Efficient Training by Optimizing Historical SolutionsTao Li, Zhehao Huang, Qinghua Tao, Yingwen Wu et al.ICLR 2023
- SUMO: Subspace-Aware Moment-Orthogonalization for Accelerating Memory-Efficient LLM TrainingYehonathan Refael, Guy Smorodinsky, Tom Tirer, Ofir LindenbaumNeurIPS 2025 · 17 citations
- Augment Your Batch: Improving Generalization Through Instance RepetitionElad Hoffer, Tal Ben-Nun, Itay Hubara, Niv Giladi et al.CVPR 2020
- Batch Normalization Orthogonalizes Representations in Deep Random NetworksHadi Daneshmand, Amir Joudaki, Francis R. BachNeurIPS 2021 · 47 citations
- RSC: Accelerate Graph Neural Networks Training via Randomized Sparse ComputationsZirui Liu, Shengyuan Chen, Kaixiong Zhou, Daochen Zha et al.ICML 2023 · 24 citations
