Threshold Differential Attention for Sink-Free, Ultra-Sparse, and Non-Dispersive Language Modeling
Xingyue Huang, Xueying Ding, Mingxuan Ju, Yozen Liu, Neil Shah, Tong Zhao
摘要
Softmax attention struggles with long contexts due to structural limitations: the strict sum-to-one constraint forces attention sinks on irrelevant tokens, and probability mass disperses as sequence lengths increase. We tackle these problems with Threshold Differential Attention (TDA), a sink-free attention mechanism that achieves ultra-sparsity and improved robustness at longer sequence lengths without the computational overhead of projection methods or the performance degradation caused by noise accumulation of standard rectified attention. TDA applies row-wise extreme-value thresholding with a length-dependent gate, retaining only exceedances. Inspired by the differential transformer, TDA also subtracts an inhibitory view to enhance expressivity. Theoretically, we prove that TDA controls the expected number of spurious survivors per row to and that consensus spurious matches across independent views vanish as context grows. Empirically, TDA produces exact zeros and eliminates attention sinks while maintaining competitive performance on standard and long-context benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 被引用 508 次
相关 Paper
- Differential TransformerTianzhu Ye, Li Dong, Yuqing Xia, Yutao Sun 等ICLR 2025
- Long-Context Generalization with Sparse AttentionPavlo Vasylenko, Hugo Pitorro, Andre F. T. Martins, Marcos V. TrevisoICLR 2026 · 被引用 19 次
- Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context FocusingLingkun Long, Yushi Huang, Shihao Bai, Ruihao Gong 等ACL 2026 · 被引用 3 次
- DSA: Efficient Inference For Video Generation Models via Distributed Sparse AttentionShenggui Li, Runyu Lu, qiaoling chen, Haiyan Yin 等ICLR 2026
- Re-ttention: Ultra Sparse Visual Generation via Attention Statistical ReshapeRuichen Chen, Keith G. Mills, Liyao Jiang, Chao Gao 等NeurIPS 2025 · 被引用 10 次
