Attention Implements the Fisher Geometry of Exponential Families
Bodie Rubacher
Abstract
Softmax attention normalizes scores, and Bayes’ rule normalizes log prior plus log likelihood. For finite latent symbols with exponential-family observations, we show that one attention head can implement the Bayes posterior and posterior means exactly, and that the posteriors representable by a single head are precisely log-linear. The standard exponential-family duality identity rewrites the likelihood as a negative Bregman divergence in mean/sufficient-statistic space; our attention-specific contribution is to use this identity to characterize when Bayes-aligned attention admits one globally shared quadratic metric, proving that this happens exactly when the dual potential is quadratic. When curvature varies, we give a multi-head local-curvature atlas with approximation and head-count bounds, and we extend the picture to in-context estimation through plug-in consistency, finite-sample stability, and an optimizer-agnostic converse from excess log-loss to approximate key-subspace alignment. Controlled Gaussian, Bernoulli, and Poisson ICE diagnostics illustrate these regimes, while the exact theorems remain scoped to finite discrete latent classes and suggest testable, not universal, predictions for larger learned transformers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 12162f91-0908-440b-a050-cdfaf2671a0cBuilds on12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong et al.NeurIPS 2023 · 356 citations
Related papers
- Statistical Advantage of Softmax Attention: Insights from Single-Location RegressionO. Duranthon, Pierre Marion, Claire Boyer, Bruno Loureiro et al.ICLR 2026 · 7 citations
- Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based PerspectiveEtienne Boursier, Claire BoyerICML 2026 · 4 citations
- In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax AttentionJianliang He, Xintian Pan, Siyu Chen, Zhuoran YangICML 2025
- In-Context Learning with Transformers: Softmax Attention Adapts to Function LipschitznessLiam Collins, Advait Parulekar, Aryan Mokhtari, Sujay Sanghavi et al.NeurIPS 2024 · 33 citations
- Dissecting the Interplay of Attention Paths in a Statistical Mechanics Theory of TransformersLorenzo Tiberi, Francesca Mignacco, Kazuki Irie, Haim SompolinskyNeurIPS 2024 · 12 citations
