TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill & Decode Inference
Xiaojuan Tang, Fanxu Meng, Pingzhi Tang, Yuxuan Wang, Di Yin, Xing Sun, Muhan Zhang
Abstract
Multi-Head Latent Attention (MLA), introduced in DeepSeek-V2, compresses key–value states into a low-rank latent vector cKV, caching only this vector to reduce memory. In tensor parallelism (TP), however, attention heads are computed across multiple devices, and each device must load the full cKV, eroding the advantage of MLA over Grouped Query Attention (GQA). We present TPLA, a scheme that partitions both the latent representation and each head's input dimension across devices, performs attention independently on each shard, and aggregates the results with an all-reduce. Unlike GLA, every attention head in TPLA still attends to the full latent space, preserving MLA's representational capacity while reducing the per-device KV cache. To make TPLA drop-in compatible with MLA checkpoints, we further derive orthogonal reparameterizations of RMSNorm and softmax---instantiated with Hadamard and PCA transforms---that mitigate cross-shard discrepancies when slicing latent vectors across devices. Finally, we introduce a prefill-decode separation scheme that keeps the MLA form during compute-bound prefilling and switches to TPLA during memory-bound decoding, minimizing conversion-induced error. By reducing the per-device KV cache for DeepSeek-V3 and Kimi-K2, we achieve 1.79x) and 1.93) speedups respectively, at a 32K-token context length while maintaining accuracy on commonsense and LongBench benchmarks. TPLA can be further implemented on top of FlashAttention-3, enabling practical end-to-end acceleration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75b035d6-c156-4047-9571-cb887dd57698Cited by top-tier papers1
Ask how each one uses itBuilds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh et al.NeurIPS 2024 · 1,019 citations
Related papers
- TyphoonMLA: A Mixed Naive-Absorb MLA Kernel For Shared PrefixAhmet Caner Yüzügüler, Ahmet Çelik, Jiawei Zhuang, Lukas CavigelliICLR 2026
- TransMLA: Migrating GQA Models to MLA with Full DeepSeek Compatibility and SpeedupFanxu Meng, Pingzhi Tang, Zengwei Yao, Xing Sun et al.NeurIPS 2025 · 5 citations
- Multi-head Temporal Latent AttentionKeqi Deng, Philip C. WoodlandNeurIPS 2025 · 2 citations
- Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMsTao Ji, Bin Guo, Yuanbin Wu, Qipeng Guo et al.ACL 2025
- Latent-Condensed Transformer for Efficient Long Context ModelingZeng You, Yaofo Chen, Qiuwu Chen, Ying Sun et al.ACL 2026
