Continuum Transformers Perform In-Context Learning by Operator Gradient Descent
Yash Patel, Abhiti Mishra, Ambuj Tewari
摘要
Transformers robustly exhibit the ability to perform in-context learning, whereby their predictive accuracy on a task can increase not by parameter updates but merely with the placement of training samples in their context windows. Recent works have shown that transformers achieve this by implementing gradient descent in their forward passes. Such results, however, are restricted to standard transformer architectures, which handle finite-dimensional inputs. In the space of PDE surrogate modeling, a generalization of transformers to handle infinite-dimensional function inputs, known as"continuum transformers,"has been proposed and similarly observed to exhibit in-context learning. Despite impressive empirical performance, such in-context learning has yet to be theoretically characterized. We herein demonstrate that continuum transformers perform in-context operator learning by performing gradient descent in an operator RKHS. We demonstrate this using novel proof strategies that leverage a generalized representer theorem for Hilbert spaces and gradient flows over the space of functionals of a Hilbert space. We additionally show the operator learned in context is the Bayes Optimal Predictor in the infinite depth limit of the transformer. We then provide empirical validations of this optimality result and demonstrate that the parameters under which such gradient descent is performed are recovered through the continuum transformer training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Fourier Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola Borislavov Kovachki, Kamyar Azizzadenesheli, Burigede Liu 等ICLR 2021 · 被引用 3,911 次
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 被引用 883 次
- Multipole Graph Neural Operator for Parametric Partial Differential EquationsZongyi Li, Nikola B. Kovachki, Kamyar Azizzadenesheli, Burigede Liu 等NeurIPS 2020 · 被引用 569 次
- Choose a Transformer: Fourier or GalerkinShuhao CaoNeurIPS 2021 · 被引用 516 次
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng 等EMNLP 2024 · 被引用 479 次
相关 Paper
- Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In ContextXiang Cheng, Yuxin Chen, Suvrit SraICML 2024 · 被引用 64 次
- What learning algorithm is in-context learning? Investigations with linear modelsEkin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma 等ICLR 2023 · 被引用 85 次
- Can Transformers Learn Full Bayesian Inference in Context?Arik Reuter, Tim G. J. Rudner, Vincent Fortuin, David RügamerICML 2025
- Transformers Can Learn Posterior Predictive Distributions In-ContextGyeonghun Kang, Changwoo Lee, Xiang ChengICML 2026
- In-context Learning on Function Classes Unveiled for TransformersZhijie Wang, Bo Jiang, Shuai LiICML 2024 · 被引用 8 次
