Attention Implements the Fisher Geometry of Exponential Families
Bodie Rubacher
摘要
Softmax attention normalizes scores, and Bayes’ rule normalizes log prior plus log likelihood. For finite latent symbols with exponential-family observations, we show that one attention head can implement the Bayes posterior and posterior means exactly, and that the posteriors representable by a single head are precisely log-linear. The standard exponential-family duality identity rewrites the likelihood as a negative Bregman divergence in mean/sufficient-statistic space; our attention-specific contribution is to use this identity to characterize when Bayes-aligned attention admits one globally shared quadratic metric, proving that this happens exactly when the dual potential is quadratic. When curvature varies, we give a multi-head local-curvature atlas with approximation and head-count bounds, and we extend the picture to in-context estimation through plug-in consistency, finite-sample stability, and an optimizer-agnostic converse from excess log-loss to approximate key-subspace alignment. Controlled Gaussian, Bernoulli, and Poisson ICE diagnostics illustrate these regimes, while the exact theorems remain scoped to finite discrete latent classes and suggest testable, not universal, predictions for larger learned transformers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 被引用 1,030 次
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 被引用 883 次
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento 等ICML 2023 · 被引用 729 次
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong 等NeurIPS 2023 · 被引用 356 次
相关 Paper
- Statistical Advantage of Softmax Attention: Insights from Single-Location RegressionO. Duranthon, Pierre Marion, Claire Boyer, Bruno Loureiro 等ICLR 2026 · 被引用 7 次
- Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based PerspectiveEtienne Boursier, Claire BoyerICML 2026 · 被引用 4 次
- In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax AttentionJianliang He, Xintian Pan, Siyu Chen, Zhuoran YangICML 2025
- In-Context Learning with Transformers: Softmax Attention Adapts to Function LipschitznessLiam Collins, Advait Parulekar, Aryan Mokhtari, Sujay Sanghavi 等NeurIPS 2024 · 被引用 33 次
- Dissecting the Interplay of Attention Paths in a Statistical Mechanics Theory of TransformersLorenzo Tiberi, Francesca Mignacco, Kazuki Irie, Haim SompolinskyNeurIPS 2024 · 被引用 12 次
