Winner-Take-All Column Row Sampling for Memory Efficient Adaptation of Language Model
Zirui Liu, Guanchu Wang, Shaochen Zhong, Zhaozhuo Xu, Daochen Zha, Ruixiang (Ryan) Tang, Zhimeng Stephen Jiang, Kaixiong Zhou, Vipin Chaudhary, Shuai Xu, Xia Hu
摘要
With the rapid growth in model size, fine-tuning the large pre-trained language model has become increasingly difficult due to its extensive memory usage. Previous works usually focus on reducing the number of trainable parameters in the network. While the model parameters do contribute to memory usage, the primary memory bottleneck during training arises from storing feature maps, also known as activations, as they are crucial for gradient calculation. Notably, neural networks are usually trained using stochastic gradient descent. We argue that in stochastic optimization, models can handle noisy gradients as long as the gradient estimator is unbiased with reasonable variance. Following this motivation, we propose a new family of unbiased estimators called WTA-CRS , for matrix production with reduced variance, which only requires storing the sub-sampled activations for calculating the gradient. Our work provides both theoretical and experimental evidence that, in the context of tuning transformers, our proposed estimators exhibit lower variance compared to existing ones. By replacing the linear operation with our approximated one in transformers, we can achieve up to 2.7× peak memory reduction with almost no accuracy drop and enables up to 6.4× larger batch size. Under the same hardware, WTA-CRS enables better down-streaming task performance by applying larger models and/or faster training speed with larger batch sizes. The code is available at https://github.com/zirui-ray-liu/WTACRS/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- One Less Reason for Filter Pruning: Gaining Free Adversarial Robustness with Structured Grouped Kernel PruningShaochen (Henry) Zhong, Zaichuan You, Jiamu Zhang, Sebastian Zhao 等NeurIPS 2023 · 被引用 13 次
- PokeMQA: Programmable knowledge editing for Multi-hop Question AnsweringHengrui Gu, Kaixiong Zhou, Xiaotian Han, Ninghao Liu 等ACL 2024 · 被引用 7 次
- SSDTrain: An Activation Offloading Framework to SSDs for Faster Large Language Model TrainingKun Wu, Jeongmin Brian Park, Xiaofan Zhang, Mert Hidayetoglu 等DAC 2025 · 被引用 3 次
- QKV Projections Require a Fraction of Their MemoryMalik Khalaf, Yara Shamshoum, Nitzan Hodos, Yuval Sieradzki 等ICLR 2026 · 被引用 3 次
- MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant OptimizationRizhen Hu, Yutong He, Ran Yan, Mou Sun 等NeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Compacter: Efficient Low-Rank Hypercomplex Adapter LayersRabeeh Karimi Mahabadi, James Henderson, Sebastian RuderNeurIPS 2021 · 被引用 700 次
- LST: Ladder Side-Tuning for Parameter and Memory Efficient Transfer LearningYi-Lin Sung, Jaemin Cho, Mohit BansalNeurIPS 2022 · 被引用 347 次
- Train Big, Then Compress: Rethinking Model Size for Efficient Training and Inference of TransformersZhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin 等ICML 2020 · 被引用 184 次
- Monarch: Expressive Structured Matrices for Efficient and Accurate TrainingTri Dao, Beidi Chen, Nimit Sharad Sohoni, Arjun D. Desai 等ICML 2022 · 被引用 125 次
相关 Paper
- PRAC: Principal-Random Subspace for LLM Activation Compression and Memory-Efficient TrainingYanyi Li, Yimu Zhang, Cong FangICML 2026
- Memory-Efficient Fine-Tuning of Transformers via Token SelectionAntoine Simoulin, Namyong Park, Xiaoyi Liu, Grey YangEMNLP 2024 · 被引用 1 次
- VeLoRA: Memory Efficient Training using Rank-1 Sub-Token ProjectionsRoy Miles, Pradyumna Reddy, Ismail Elezi, Jiankang DengNeurIPS 2024 · 被引用 22 次
- Softmax Output Approximation for Activation Memory-Efficient Training of Attention-based NetworksChanghyeon Lee, Seulki LeeNeurIPS 2023 · 被引用 4 次
- Tempo: Accelerating Transformer-Based Model Training through Memory Footprint ReductionMuralidhar Andoorveedu, Zhanda Zhu, Bojian Zheng, Gennady PekhimenkoNeurIPS 2022 · 被引用 8 次
