Winner-Take-All Column Row Sampling for Memory Efficient Adaptation of Language Model
Zirui Liu, Guanchu Wang, Shaochen Zhong, Zhaozhuo Xu, Daochen Zha, Ruixiang (Ryan) Tang, Zhimeng Stephen Jiang, Kaixiong Zhou, Vipin Chaudhary, Shuai Xu, Xia Hu
Abstract
With the rapid growth in model size, fine-tuning the large pre-trained language model has become increasingly difficult due to its extensive memory usage. Previous works usually focus on reducing the number of trainable parameters in the network. While the model parameters do contribute to memory usage, the primary memory bottleneck during training arises from storing feature maps, also known as activations, as they are crucial for gradient calculation. Notably, neural networks are usually trained using stochastic gradient descent. We argue that in stochastic optimization, models can handle noisy gradients as long as the gradient estimator is unbiased with reasonable variance. Following this motivation, we propose a new family of unbiased estimators called WTA-CRS , for matrix production with reduced variance, which only requires storing the sub-sampled activations for calculating the gradient. Our work provides both theoretical and experimental evidence that, in the context of tuning transformers, our proposed estimators exhibit lower variance compared to existing ones. By replacing the linear operation with our approximated one in transformers, we can achieve up to 2.7× peak memory reduction with almost no accuracy drop and enables up to 6.4× larger batch size. Under the same hardware, WTA-CRS enables better down-streaming task performance by applying larger models and/or faster training speed with larger batch sizes. The code is available at https://github.com/zirui-ray-liu/WTACRS/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3c3dd916-d0ae-4860-8a5b-d535363fec45Cited by top-tier papers8
- One Less Reason for Filter Pruning: Gaining Free Adversarial Robustness with Structured Grouped Kernel PruningShaochen (Henry) Zhong, Zaichuan You, Jiamu Zhang, Sebastian Zhao et al.NeurIPS 2023 · 13 citations
- PokeMQA: Programmable knowledge editing for Multi-hop Question AnsweringHengrui Gu, Kaixiong Zhou, Xiaotian Han, Ninghao Liu et al.ACL 2024 · 7 citations
- SSDTrain: An Activation Offloading Framework to SSDs for Faster Large Language Model TrainingKun Wu, Jeongmin Brian Park, Xiaofan Zhang, Mert Hidayetoglu et al.DAC 2025 · 3 citations
- QKV Projections Require a Fraction of Their MemoryMalik Khalaf, Yara Shamshoum, Nitzan Hodos, Yuval Sieradzki et al.ICLR 2026 · 3 citations
- MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant OptimizationRizhen Hu, Yutong He, Ran Yan, Mou Sun et al.NeurIPS 2025 · 1 citation
Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Compacter: Efficient Low-Rank Hypercomplex Adapter LayersRabeeh Karimi Mahabadi, James Henderson, Sebastian RuderNeurIPS 2021 · 700 citations
- LST: Ladder Side-Tuning for Parameter and Memory Efficient Transfer LearningYi-Lin Sung, Jaemin Cho, Mohit BansalNeurIPS 2022 · 347 citations
- Train Big, Then Compress: Rethinking Model Size for Efficient Training and Inference of TransformersZhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin et al.ICML 2020 · 184 citations
- Monarch: Expressive Structured Matrices for Efficient and Accurate TrainingTri Dao, Beidi Chen, Nimit Sharad Sohoni, Arjun D. Desai et al.ICML 2022 · 125 citations
Related papers
- PRAC: Principal-Random Subspace for LLM Activation Compression and Memory-Efficient TrainingYanyi Li, Yimu Zhang, Cong FangICML 2026
- Memory-Efficient Fine-Tuning of Transformers via Token SelectionAntoine Simoulin, Namyong Park, Xiaoyi Liu, Grey YangEMNLP 2024 · 1 citation
- VeLoRA: Memory Efficient Training using Rank-1 Sub-Token ProjectionsRoy Miles, Pradyumna Reddy, Ismail Elezi, Jiankang DengNeurIPS 2024 · 22 citations
- Softmax Output Approximation for Activation Memory-Efficient Training of Attention-based NetworksChanghyeon Lee, Seulki LeeNeurIPS 2023 · 4 citations
- Tempo: Accelerating Transformer-Based Model Training through Memory Footprint ReductionMuralidhar Andoorveedu, Zhanda Zhu, Bojian Zheng, Gennady PekhimenkoNeurIPS 2022 · 8 citations
