Tensor Product Attention Is All You Need
Yifan Zhang, Yifeng Liu, Huizhuo Yuan, Zhen Qin, Yang Yuan, Quanquan Gu, Andrew C. Yao
摘要
Scaling language models to handle longer input sequences typically necessitates large key-value (KV) caches, resulting in substantial memory overhead during inference. In this paper, we propose Tensor Product Attention (TPA), a novel attention mechanism that uses tensor decompositions to represent queries, keys, and values compactly, substantially shrinking the KV cache size at inference time. By factorizing these representations into contextual low-rank components and seamlessly integrating with Rotary Position Embedding (RoPE), TPA achieves improved model quality alongside memory efficiency. Based on TPA, we introduce the Tensor ProducT ATTenTion Transformer (T6), a new model architecture for sequence modeling. Through extensive empirical evaluation on language modeling tasks, we demonstrate that T6 surpasses or matches the performance of standard Transformer baselines including Multi-Head Attention (MHA), Multi-Query Attention (MQA), Grouped-Query Attention (GQA), and Multi-Head Latent Attention (MLA) across various metrics, including perplexity and a range of established evaluation benchmarks. Notably, TPA's memory efficiency and computational efficiency at decoding stage enables processing longer sequences under fixed resource constraints, addressing a critical scalability challenge in modern language models. Project Page: https://github.com/tensorgi/TPA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Cartridges: Lightweight and general-purpose long context representations via self-studySabri Eyuboglu, Ryan Ehrlich, Simran Arora, Neel Guha 等ICLR 2026 · 被引用 74 次
- Multi-Head Low-Rank AttentionSongtao Liu, Hongwu Peng, Zhiwei Zhang, Zhengyu Chen 等ICLR 2026 · 被引用 18 次
- MUSTAFAR: Promoting Unstructured Sparsity for KV Cache Pruning in LLM InferenceDonghyeon Joo, Helya Hosseini, Ramyad Hadidi, Bahar AsgariNeurIPS 2025 · 被引用 12 次
- UniGist: Towards General and Hardware-aligned Sequence-level Long Context CompressionChenlong Deng, Zhisong Zhang, Kelong Mao, Shuaiyi Li 等NeurIPS 2025 · 被引用 10 次
- TransMLA: Migrating GQA Models to MLA with Full DeepSeek Compatibility and SpeedupFanxu Meng, Pingzhi Tang, Zengwei Yao, Xing Sun 等NeurIPS 2025 · 被引用 5 次
它引用的顶会 Paper29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
相关 Paper
- SALS: Sparse Attention in Latent Space for KV Cache CompressionJunlin Mu, Hantao Huang, Jihang Zhang, Minghui Yu 等NeurIPS 2025 · 被引用 7 次
- KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional EmbeddingLuohe Shi, Zuchao Li, Lefei Zhang, Baoyuan Qi 等ACL 2025 · 被引用 5 次
- Tucker Attention: A generalization of approximate attention mechanismsTimon Klein, Jonas Kusch, Sebastian Sager, Stefan Schnake 等ICML 2026 · 被引用 3 次
- SWAN: An Efficient and Scalable Approach for Long-Context Language ModelingKrishna C. Puvvada, Faisal Ladhak, Santiago Akle Serano, Cheng-Ping Hsieh 等EMNLP 2025
- CommVQ: Commutative Vector Quantization for KV Cache CompressionJunyan Li, Yang Zhang, Muhammad Yusuf Hassan, Talha Chafekar 等ICML 2025
