DINT Transformer
Yueyang Cang, Yuhang Liu, Xiaoteng Zhang, Erlu Zhao, Li Shi
摘要
The DIFF Transformer mitigates interference from irrelevant contexts by introducing a differential attention mechanism, thereby enhancing focus on critical tokens. However, this architecture suffers from two major limitations: first, its use of two independent attention matrices leads to numerical instability, and second, it lacks global context modeling, which is essential for identifying globally significant tokens. To address these challenges, we propose the DINT Transformer, which extends the DIFF Transformer by incorporating an integral mechanism. By computing global importance scores and integrating them into the attention matrix, the DINT Transformer not only improves overall numerical stability but also significantly enhances its ability to capture global dependencies. Experimental results demonstrate that the DINT Transformer achieves superior accuracy and robustness across various practical applications, including long-context language modeling and key information retrieval. These advancements establish the DINT Transformer as a highly effective and promising architecture.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel 等ACL 2022 · 被引用 1,494 次
- The Devil in Linear TransformerZhen Qin, Xiaodong Han, Weixuan Sun, Dongxu Li 等EMNLP 2022 · 被引用 24 次
- Differential TransformerTianzhu Ye, Li Dong, Yuqing Xia, Yutao Sun 等ICLR 2025
相关 Paper
- Integral Transformer: Denoising Attention, Not Too Much Not Too LittleIvan Kobyzev, Abbas Ghaddar, Dingtao Hu, Boxing ChenEMNLP 2025
- LUCID: Attention with Preconditioned RepresentationsSai Surya Duvvuri, Nirmal Patel, Nilesh Gupta, Inderjit DhillonICML 2026
- Inceptive Transformers: Enhancing Contextual Representations through Multi-Scale Feature Learning Across Domains and LanguagesAsif Shahriar, Rifat Shahriyar, M. Saifur RahmanEMNLP 2025
- Towards Stabilized and Efficient Diffusion Transformers Through Long-Skip-Connections With Spectral ConstraintsGuanjie Chen, Xinyu Zhao, Yucheng Zhou, Xiaoye Qu 等ICCV 2025 · 被引用 1 次
- Dynamic Linear AttentionXin Wang, Hui Shen, Boyuan Zheng, Xueshen Liu 等ICML 2026 · 被引用 1 次
