QKV Projections Require a Fraction of Their Memory
Malik Khalaf, Yara Shamshoum, Nitzan Hodos, Yuval Sieradzki, Assaf Schuster
Abstract
The Multi-Head Attention mechanism is central to LLM operation, and multiple works target its compute and memory efficiency during training. While most works focus on approximating the scaled dot product, the memory consumption of the linear projections that compute the , , and tensors from the input is often overlooked. To address this, we propose Point-Approximate Matrix Multiplication (PAMM), a novel tensor compression technique that compresses the activations of the projections in attention layers by a factor of up to , effectively erasing their memory footprint, while achieving similar or better final perplexity. PAMM is fully composable with efficient attention techniques such as FlashAttention, making it a practical and complementary method for memory-efficient LLM training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e90d4689-5350-4df2-8a00-6021ebd8f0caBuilds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
Related papers
- Tucker Attention: A generalization of approximate attention mechanismsTimon Klein, Jonas Kusch, Sebastian Sager, Stefan Schnake et al.ICML 2026 · 3 citations
- ESPACE: Dimensionality Reduction of Activations for Model CompressionCharbel Sakr, Brucek KhailanyNeurIPS 2024 · 21 citations
- AQPIM: Breaking the PIM Capacity Wall for LLMs with in-Memory Activation QuantizationKosuke Matsushima, Yasuyuki Okoshi, Masato Motomura, Daichi FujikiHPCA 2026 · 1 citation
- Low-Rank Approximation for Sparse Attention in Multi-Modal LLMsLin Song, Yukang Chen, Shuai Yang, Xiaohan Ding et al.CVPR 2024 · 8 citations
- Palu: KV-Cache Compression with Low-Rank ProjectionChi-Chih Chang, Wei-Cheng Lin, Chien-Yu Lin, Chong-Yan Chen et al.ICLR 2025
