Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence Transformers
Yukun Zhang, Xueqing Zhou
摘要
We present Continuous-Time Attention, a novel framework that infuses partial differential equations (PDEs) into the Transformer's attention mechanism to better handle long sequences. Instead of relying on a static attention matrix, we allow attention weights to evolve along a pseudo-time dimension governed by diffusion, wave, or reaction-diffusion dynamics. This dynamic process systematically smooths local noise, strengthens long-range dependencies, and improves gradient stability during training.Our theoretical analysis shows that PDE-driven attention mitigates the exponential decay of distant interactions and improves the optimization landscape. Empirically, Continuous-Time Attention achieves consistent performance gains over both standard and long-sequence Transformer variants across a range of tasks. These results suggest that embedding continuous-time dynamics into attention mechanisms is a promising direction for enhancing global coherence and scalability in Transformer models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier 等ICLR 2020 · 被引用 833 次
- Learning Differential Equations that are Easy to SolveJacob Kelly, Jesse Bettencourt, Matthew J. Johnson, David DuvenaudNeurIPS 2020 · 被引用 134 次
相关 Paper
- ContiFormer: Continuous-Time Transformer for Irregular Time Series ModelingYuqi Chen, Kan Ren, Yansen Wang, Yuchen Fang 等NeurIPS 2023 · 被引用 131 次
- Wavy TransformerSatoshi Noguchi, Yoshinobu KawaharaNeurIPS 2025 · 被引用 1 次
- Robust Filter Attention: Self-Attention as a Parallel State EstimatorPeter RacioppoICML 2026
- Rough Transformers: Lightweight and Continuous Time Series Modelling through Signature PatchingFernando Moreno-Pino, Alvaro Arroyo, Harrison Waldon, Xiaowen Dong 等NeurIPS 2024 · 被引用 21 次
- AROMA: Preserving Spatial Structure for Latent PDE Modeling with Local Neural FieldsLouis Serrano, Thomas X. Wang, Etienne Le Naour, Jean-Noël Vittaut 等NeurIPS 2024 · 被引用 47 次
