Differential Transformer
Tianzhu Ye, Li Dong, Yuqing Xia, Yutao Sun, Yi Zhu, Gao Huang, Furu Wei
摘要
Transformer tends to overallocate attention to irrelevant context. In this work, we introduce DIFF Transformer, which amplifies attention to the relevant context while canceling noise. Specifically, the differential attention mechanism calculates attention scores as the difference between two separate softmax attention maps. The subtraction cancels noise, promoting the emergence of sparse attention patterns. Experimental results on language modeling show that DIFF Transformer outperforms Transformer in various settings of scaling up model size and training tokens. More intriguingly, it offers notable advantages in practical applications, such as long-context modeling, key information retrieval, hallucination mitigation, in-context learning, and reduction of activation outliers. By being less distracted by irrelevant context, DIFF Transformer can mitigate hallucination in question answering and text summarization. For in-context learning, DIFF Transformer not only enhances accuracy but is also more robust to order permutation, which was considered as a chronic robustness issue. The results position DIFF Transformer as a highly effective and promising architecture to advance large language models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper41
- Neural Attention SearchDifan Deng, Marius LindauerNeurIPS 2025 · 被引用 431 次
- Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-FreeZihan Qiu, Zekun Wang, Bo Zheng, Zeyu Huang 等NeurIPS 2025 · 被引用 336 次
- Rope to Nope and Back Again: A New Hybrid Attention StrategyBowen Yang, Bharat Venkitesh, Dwaraknath Gnaneshwar, Hangyu Lin 等NeurIPS 2025 · 被引用 51 次
- Doc-to-LoRA: Learning to Instantly Internalize ContextsRujikorn Charakorn, Edoardo Cetin, Shinnosuke Uesaka, Robert LangeICML 2026 · 被引用 27 次
- scDFM: Distributional Flow Matching Model for Robust Single-Cell Perturbation PredictionChenglei Yu, Chuanrui Wang, Bangyan Liao, Tailin WuICLR 2026 · 被引用 15 次
它引用的顶会 Paper16
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel 等ACL 2022 · 被引用 1,494 次
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionJay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar 等NeurIPS 2024 · 被引用 727 次
- Data Engineering for Scaling Language Models to 128K ContextYao Fu, Rameswar Panda, Xinyao Niu, Xiang Yue 等ICML 2024 · 被引用 204 次
相关 Paper
- DINT TransformerYueyang Cang, Yuhang Liu, Xiaoteng Zhang, Erlu Zhao 等EMNLP 2025 · 被引用 1 次
- Understanding Differential Transformer Unchains Pretrained Self-AttentionsChaerin Kong, Jiho Jang, Nojun KwakNeurIPS 2025
- Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language ModelingXingyue Huang, Xueying Ding, Mingxuan Ju, Yozen Liu 等ACL 2026 · 被引用 3 次
- Efficient OpAmp Adaptation for Zoom Attention to Golden ContextsHaoyuan Wu, Rui Ming, Haisheng Zheng, Zhuolun He 等ACL 2025 · 被引用 1 次
- Integral Transformer: Denoising Attention, Not Too Much Not Too LittleIvan Kobyzev, Abbas Ghaddar, Dingtao Hu, Boxing ChenEMNLP 2025
