Kalman Linear Attention: Parallel Bayesian Filtering For Efficient Language Modeling and State Tracking
Vaisakh Shaj, Cameron Barker, Aidan Scannell, Andras Szecsenyi, Elliot Crowley, Amos Storkey
Abstract
State-space language models such as Mamba and gated linear attention (GLA) offer linearcomplexity, parallelisable alternatives to transformers, but their linear state updates limit expressivity and robust state tracking. We close this gap from a probabilistic angle, casting sequence mixing as exact Bayesian filtering with the Kalman filter as the core primitive. Classical Kalman filters give principled state and uncertainty estimates but are viewed as inherently sequential; we show that reparameterising them in information form turns their updates into an associative scan -so the per-token recurrent update is non-linear (a Möbius/precision recursion) yet remains temporally parallel. The resulting Kalman Linear Attention (KLA) layer is a drop-in sequence mixer that performs time-parallel probabilistic inference, carries an explicit belief-state uncertainty, and is strictly more expressive than GLA-style linear updates at the same computational cost. This expressivity translates directly into stronger state tracking: KLA solves permutation-composition (A 5 ) tasks that linear SSMs and attention cannot, while staying scan-parallel. As a drop-in primitive it also matches or improves on modern SSMs and GLAs across synthetic token-manipulation and zero-shot commonsense benchmarks, and is among the first stacked Bayesian-filtering primitives trained at the billion-token scale. We introduce Kalman Linear Attention (KLA), which formulates sequence modelling as a Bayesian filtering problem. KLA models two sources of uncertainty: process noise, which captures uncertainty in state evolution, and
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6504eff7-3fcd-4222-95c5-3d64502ebf3aBuilds on19
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
Related papers
- Sequential Parallel Duality in Prefix Scannable ModelsMorris Yau, Sharut Gupta, Valerie Engelmayer, Kazuki Irie et al.ICLR 2026 · 9 citations
- Fixed-Point RNNs: Interpolating from Diagonal to DenseSajad Movahedi, Felix Sarnthein, Nicola Muca Cirone, Antonio OrvietoNeurIPS 2025 · 5 citations
- Gated KalmaNet: A Fading Memory Layer through Test-time Ridge RegressionLiangzu Peng, Aditya Chattopadhyay, Luca Zancato, Elvis Nunez et al.CVPR 2026 · 10 citations
- Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language ModelingLiliang Ren, Yang Liu, Yadong Lu, Yelong Shen et al.ICLR 2025
- The Expressive Capacity of State Space Models: A Formal Language PerspectiveYash Raj Sarrof, Yana Veitsman, Michael HahnNeurIPS 2024 · 53 citations
