Improving Transformers with Dynamically Composable Multi-Head Attention
Da Xiao, Qingye Meng, Shengping Li, Xingyuan Yuan
摘要
Multi-Head Attention (MHA) is a key component of Transformer. In MHA, attention heads work independently, causing problems such as low-rank bottleneck of attention score matrices and head redundancy. We propose Dynamically Composable Multi-Head Attention (DCMHA), a parameter and computation efficient attention architecture that tackles the shortcomings of MHA and increases the expressive power of the model by dynamically composing attention heads. At the core of DCMHA is a function that transforms the attention score and weight matrices in an input-dependent way. DCMHA can be used as a drop-in replacement of MHA in any transformer architecture to obtain the corresponding DCFormer. DCFormer significantly outperforms Transformer on different architectures and model scales in language modeling, matching the performance of models with 1.7x-2.0x compute. For example, DCPythia-6.9B outperforms open source Pythia-12B on both pretraining perplexity and downstream task evaluation. The code and models are available at https://github.com/Caiyun-AI/DCFormer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-FreeZihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang 等NeurIPS 2025 · 被引用 336 次
- DAPE V2: Process Attention Score as Feature Map for Length ExtrapolationChuanyang Zheng, Yihang Gao, Han Shi, Jing Xiong 等ACL 2025 · 被引用 12 次
- Language Modeling by Language ModelsJunyan Cheng, Peter Clark, Kyle RichardsonNeurIPS 2025 · 被引用 11 次
- Efficient Degradation-agnostic Image Restoration via Channel-Wise Functional Decomposition and Manifold RegularizationBin Ren, Yawei Li, Xu Zheng, Yuqian Fu 等ICLR 2026 · 被引用 9 次
- Devil is in the Uniformity: Exploring Diverse Learners Within Transformer for Image RestorationShihao Zhou, Dayu Li, Jinshan Pan, Juncheng Zhou 等ICCV 2025 · 被引用 8 次
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
相关 Paper
- MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense ConnectionsDa Xiao, Qingye Meng, Shengping Li, Xingyuan YuanICML 2025
- SAS: Simulated Attention ScoreChuanyang Zheng, Jiankai Sun, Yihang Gao, Yuehao Wang 等NeurIPS 2025 · 被引用 3 次
- DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads FusionYilong Chen, Linhao Zhang, Junyuan Shang, Zhenyu Zhang 等NeurIPS 2024 · 被引用 12 次
- MoH: Multi-Head Attention as Mixture-of-Head AttentionPeng Jin, Bo Zhu, Li Yuan, Shuicheng YanICML 2025 · 被引用 2 次
- Mixture of Attention Heads: Selecting Attention Heads Per TokenXiaofeng Zhang, Yikang Shen, Zeyu Huang, Jie Zhou 等EMNLP 2022 · 被引用 23 次
