Tensor Product Attention Is All You Need
Yifan Zhang, Yifeng Liu, Huizhuo Yuan, Zhen Qin, Yang Yuan, Quanquan Gu, Andrew C. Yao
Abstract
Scaling language models to handle longer input sequences typically necessitates large key-value (KV) caches, resulting in substantial memory overhead during inference. In this paper, we propose Tensor Product Attention (TPA), a novel attention mechanism that uses tensor decompositions to represent queries, keys, and values compactly, substantially shrinking the KV cache size at inference time. By factorizing these representations into contextual low-rank components and seamlessly integrating with Rotary Position Embedding (RoPE), TPA achieves improved model quality alongside memory efficiency. Based on TPA, we introduce the Tensor ProducT ATTenTion Transformer (T6), a new model architecture for sequence modeling. Through extensive empirical evaluation on language modeling tasks, we demonstrate that T6 surpasses or matches the performance of standard Transformer baselines including Multi-Head Attention (MHA), Multi-Query Attention (MQA), Grouped-Query Attention (GQA), and Multi-Head Latent Attention (MLA) across various metrics, including perplexity and a range of established evaluation benchmarks. Notably, TPA's memory efficiency and computational efficiency at decoding stage enables processing longer sequences under fixed resource constraints, addressing a critical scalability challenge in modern language models. Project Page: https://github.com/tensorgi/TPA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2515991-0782-41e8-bc7f-1f040ec63054Cited by top-tier papers11
- Cartridges: Lightweight and general-purpose long context representations via self-studySabri Eyuboglu, Ryan Ehrlich, Simran Arora, Neel Guha et al.ICLR 2026 · 74 citations
- Multi-Head Low-Rank AttentionSongtao Liu, Hongwu Peng, Zhiwei Zhang, Zhengyu Chen et al.ICLR 2026 · 18 citations
- MUSTAFAR: Promoting Unstructured Sparsity for KV Cache Pruning in LLM InferenceDonghyeon Joo, Helya Hosseini, Ramyad Hadidi, Bahar AsgariNeurIPS 2025 · 12 citations
- UniGist: Towards General and Hardware-aligned Sequence-level Long Context CompressionChenlong Deng, Zhisong Zhang, Kelong Mao, Shuaiyi Li et al.NeurIPS 2025 · 10 citations
- TransMLA: Migrating GQA Models to MLA with Full DeepSeek Compatibility and SpeedupFanxu Meng, Pingzhi Tang, Zengwei Yao, Xing Sun et al.NeurIPS 2025 · 5 citations
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
Related papers
- SALS: Sparse Attention in Latent Space for KV Cache CompressionJunlin Mu, Hantao Huang, Jihang Zhang, Minghui Yu et al.NeurIPS 2025 · 7 citations
- KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional EmbeddingLuohe Shi, Zuchao Li, Lefei Zhang, Baoyuan Qi et al.ACL 2025 · 5 citations
- Tucker Attention: A generalization of approximate attention mechanismsTimon Klein, Jonas Kusch, Sebastian Sager, Stefan Schnake et al.ICML 2026 · 3 citations
- SWAN: An Efficient and Scalable Approach for Long-Context Language ModelingKrishna C. Puvvada, Faisal Ladhak, Santiago Akle Serano, Cheng-Ping Hsieh et al.EMNLP 2025
- CommVQ: Commutative Vector Quantization for KV Cache CompressionJunyan Li, Yang Zhang, Muhammad Yusuf Hassan, Talha Chafekar et al.ICML 2025
