Efficient Backpropagation with Variance Controlled Adaptive Sampling
Ziteng Wang, Jianfei Chen, Jun Zhu
Abstract
Sampling-based algorithms, which eliminate ''unimportant'' computations during forward and/or back propagation (BP), offer potential solutions to accelerate neural network training. However, since sampling introduces approximations to training, such algorithms may not consistently maintain accuracy across various tasks. In this work, we introduce a variance-controlled adaptive sampling (VCAS) method designed to accelerate BP. VCAS computes an unbiased stochastic gradient with fine-grained layerwise importance sampling in data dimension for activation gradient calculation and leverage score sampling in token dimension for weight gradient calculation. To preserve accuracy, we control the additional variance by learning the sample ratio jointly with model parameters during training. We assessed VCAS on multiple fine-tuning and pre-training tasks in both vision and natural language domains. On all the tasks, VCAS can preserve the original training loss trajectory and validation accuracy with an up to 73.87% FLOPs reduction of BP and 49.58% FLOPs reduction of the whole training process. The implementation is available at https://github.com/thu-ml/VCAS .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d13cae5f-c987-41b8-8a6c-5b05171cf716Cited by top-tier papers2
- SAGE: A Framework of Precise Retrieval for RAGJintao Zhang, Guoliang Li, Jinyang SuICDE 2025 · 9 citations
- STAFF: Speculative Coreset Selection for Task-Specific Fine-tuningXiaoyu Zhang, Juan Zhai, Shiqing Ma, Chao Shen et al.ICLR 2025
Builds on11
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman et al.ICLR 2020 · 462 citations
- Prioritized Training on Points that are Learnable, Worth Learning, and not yet LearntSören Mindermann, Jan Markus Brauner, Muhammed Razzak, Mrinank Sharma et al.ICML 2022 · 237 citations
- The Efficiency MisnomerMostafa Dehghani, Yi Tay, Anurag Arnab, Lucas Beyer et al.ICLR 2022 · 116 citations
- Cramming: Training a Language Model on a single GPU in one dayJonas Geiping, Tom GoldsteinICML 2023 · 115 citations
Related papers
- One Forward is Enough for Neural Network Training via Likelihood Ratio MethodJinyang Jiang, Zeliang Zhang, Chenliang Xu, Zhaofei Yu et al.ICLR 2024 · 14 citations
- ADA-GP: Accelerating DNN Training By Adaptive Gradient PredictionVahid Janfaza, Shantanu Mandal, Farabi Mahmud, Abdullah MuzahidMICRO 2023 · 3 citations
- MISA: Memory-Efficient LLMs Optimization with Module-wise Importance SamplingYuxi Liu, Renjia Deng, Yutong He, Xue Wang et al.NeurIPS 2025 · 2 citations
- Winner-Take-All Column Row Sampling for Memory Efficient Adaptation of Language ModelZirui Liu, Guanchu Wang, Shaochen Zhong, Zhaozhuo Xu et al.NeurIPS 2023 · 22 citations
- Token Dropping for Efficient BERT PretrainingLe Hou, Richard Yuanzhe Pang, Tianyi Zhou, Yuexin Wu et al.ACL 2022
