Lune

ICLR2026Top-tier venue

Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality

Ryotaro Kawata, Taiji Suzuki

2026Year
3Citations

Abstract

Transformers excel through content-addressable retrieval and the ability to exploit contexts of, in principle, unbounded length. We recast associative memory at the level of probability measures, treating a context as a distribution over tokens and viewing attention as an integral operator on measures. Concretely, for mixture contexts ν=I−1∑i=1Iμ(i)\nu = I^{-1} \sum_{i=1}^I \mu^{(i)} and a query xq(i\*)x_{\mathrm{q}}(i^\*), the task decomposes into (i) recall of the relevant component μ(i\*)\mu^{(i^\*)} and (ii) prediction from (μi\*,xq)(\mu_{i^\*},x_{\mathrm{q}}). We study learned softmax attention (not a frozen kernel) trained by empirical risk minimization and show that a shallow measure-theoretic Transformer composed with an MLP learns the recall-and-predict map under a spectral assumption on the input densities. We further establish a matching minimax lower bound with the same rate exponent (up to multiplicative constants), proving sharpness of the convergence order. The framework offers a principled recipe for designing and analyzing Transformers that recall from arbitrarily long, distributional contexts with provable generalization guarantees.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext bca89964-5bc3-41fe-a569-edd03a616f59

Builds on21

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines