DropIT: Dropping Intermediate Tensors for Memory-Efficient DNN Training
Joya Chen, Kai Xu, Yuhui Wang, Yifei Cheng, Angela Yao
摘要
A standard hardware bottleneck when training deep neural networks is GPU memory. The bulk of memory is occupied by caching intermediate tensors for gradient computation in the backward pass. We propose a novel method to reduce this footprint - Dropping Intermediate Tensors (DropIT). DropIT drops min-k elements of the intermediate tensors and approximates gradients from the sparsified tensors in the backward pass. Theoretically, DropIT reduces noise on estimated gradients and therefore has a higher rate of convergence than vanilla-SGD. Experiments show that we can drop up to 90% of the intermediate tensor elements in fully-connected and convolutional layers while achieving higher testing accuracy for Visual Transformers and Convolutional Neural Networks on various tasks (e.g., classification, object detection, instance segmentation). Our code and models are available at https://github.com/chenjoya/dropit.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Scaling for Training Time and Post-hoc Out-of-distribution Detection EnhancementKai Xu, Rongyu Chen, Gianni Franchi, Angela YaoICLR 2024 · 被引用 81 次
- An Efficient and Accurate Dynamic Sparse Training Framework Based on Parameter-FreezingLei Li, Haochen Yang, Jiacheng Guo, Hongkai Yu 等AAAI 2025 · 被引用 2 次
- SURGEON: Memory-Adaptive Fully Test-Time Adaptation via Dynamic Activation SparsityKe Ma, Jiaqi Tang, Bin Guo, Fan Dang 等CVPR 2025
- PaCA: Partial Connection Adaptation for Efficient Fine-TuningSunghyeon Woo, Sol Namkung, Sunwoo Lee, Inho Jeong 等ICLR 2025
- MIST : Multi-modal Iterative Spatial-Temporal Transformer for Long-form Video Question AnsweringDifei Gao, Luowei Zhou, Lei Ji, Linchao Zhu 等CVPR 2023
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign DropoutZhao Chen, Jiquan Ngiam, Yanping Huang, Thang Luong 等NeurIPS 2020 · 被引用 313 次
相关 Paper
- An In-depth Study of Stochastic BackpropagationJun Fang, Mingze Xu, Hao Chen, Bing Shuai 等NeurIPS 2022 · 被引用 2 次
- Stochastic Backpropagation: A Memory Efficient Strategy for Training Video ModelsFeng Cheng, Mingze Xu, Yuanjun Xiong, Hao Chen 等CVPR 2022 · 被引用 11 次
- Optimal Gradient Checkpoint Search for Arbitrary Computation GraphsJianwei Feng, Dong HuangCVPR 2021
- Memory Optimization for Deep NetworksAashaka Shah, Chao-Yuan Wu, Jayashree Mohan, Vijay Chidambaram 等ICLR 2021 · 被引用 29 次
- INSTANT: Compressing Gradients and Activations for Resource-Efficient TrainingTuan-Kiet Doan, Trung-Hieu Tran, Enzo Tartaglione, Nikola Simidjievski 等ICLR 2026
