DINT Transformer
Yueyang Cang, Yuhang Liu, Xiaoteng Zhang, Erlu Zhao, Li Shi
Abstract
The DIFF Transformer mitigates interference from irrelevant contexts by introducing a differential attention mechanism, thereby enhancing focus on critical tokens. However, this architecture suffers from two major limitations: first, its use of two independent attention matrices leads to numerical instability, and second, it lacks global context modeling, which is essential for identifying globally significant tokens. To address these challenges, we propose the DINT Transformer, which extends the DIFF Transformer by incorporating an integral mechanism. By computing global importance scores and integrating them into the attention matrix, the DINT Transformer not only improves overall numerical stability but also significantly enhances its ability to capture global dependencies. Experimental results demonstrate that the DINT Transformer achieves superior accuracy and robustness across various practical applications, including long-context language modeling and key information retrieval. These advancements establish the DINT Transformer as a highly effective and promising architecture.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 375ec4a5-4aca-41cc-86f8-c8b01a0cd428Cited by top-tier papers1
Ask how each one uses itBuilds on3
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
- The Devil in Linear TransformerZhen Qin, Xiaodong Han, Weixuan Sun, Dongxu Li et al.EMNLP 2022 · 24 citations
- Differential TransformerTianzhu Ye, Li Dong, Yuqing Xia, Yutao Sun et al.ICLR 2025
Related papers
- Integral Transformer: Denoising Attention, Not Too Much Not Too LittleIvan Kobyzev, Abbas Ghaddar, Dingtao Hu, Boxing ChenEMNLP 2025
- LUCID: Attention with Preconditioned RepresentationsSai Surya Duvvuri, Nirmal Patel, Nilesh Gupta, Inderjit DhillonICML 2026
- Inceptive Transformers: Enhancing Contextual Representations through Multi-Scale Feature Learning Across Domains and LanguagesAsif Shahriar, Rifat Shahriyar, M. Saifur RahmanEMNLP 2025
- Towards Stabilized and Efficient Diffusion Transformers Through Long-Skip-Connections With Spectral ConstraintsGuanjie Chen, Xinyu Zhao, Yucheng Zhou, Xiaoye Qu et al.ICCV 2025 · 1 citation
- Dynamic Linear AttentionXin Wang, Hui Shen, Boyuan Zheng, Xueshen Liu et al.ICML 2026 · 1 citation
