Robust Filter Attention: Self-Attention as a Parallel State Estimator
Peter Racioppo
Abstract
We introduce Robust Filter Attention (RFA), an attention mechanism that reformulates self-attention as parallel robust filtering under a latent stochastic differential equation (SDE) prior, where analytically propagated uncertainty defines a time-dependent precision prior over attention weights. This formulation integrates key advantages of existing positional encodings: it preserves RoPE-style rotational structure while achieving long-context stability through explicit modeling of dissipation and diffusion. By imposing isotropic constraints on the dynamics and noise, RFA matches the time and memory complexity of standard attention. Empirically, we find that uncertainty-aware weighting induces specialization into distinct filtering regimes across heads, improving temporal consistency and extrapolation across varying context lengths.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on13
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- HiPPO: Recurrent Memory with Optimal Polynomial ProjectionsAlbert Gu, Tri Dao, Stefano Ermon, Atri Rudra et al.NeurIPS 2020 · 1,100 citations
- Hopfield Networks is All You NeedHubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl et al.ICLR 2021 · 620 citations
Related papers
- Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence TransformersYukun Zhang, Xueqing ZhouEMNLP 2025
- LoL: Longer than Longer, Scaling Video Generation to HourJustin Cui, Jie Wu, Ming Li, Tao Yang et al.CVPR 2026 · 30 citations
- Wavelet-based Positional Representation for Long ContextYui Oka, Taku Hasegawa, Kyosuke Nishida, Kuniko SaitoICLR 2025
- Beyond Position: the emergence of wavelet-like properties in TransformersValeria Ruscio, Umberto Nanni, Fabrizio SilvestriACL 2025 · 1 citation
- SWAN: An Efficient and Scalable Approach for Long-Context Language ModelingKrishna C. Puvvada, Faisal Ladhak, Santiago Akle Serano, Cheng-Ping Hsieh et al.EMNLP 2025
