Nonlinear transformers can perform inference-time feature learning
Naoki Nishikawa, Yujin Song, Kazusato Oko, Denny Wu, Taiji Suzuki
摘要
Pretrained transformers have demonstrated the ability to implement various algorithms at inference time without parameter updates. While theoretical works have established this capability through constructions and approximation guarantees, the optimization and statistical efficiency aspects remain understudied. In this work, we investigate how transformers learn features in-contexta key mechanism underlying their inference-time adaptivity. We focus on the in-context learning of single-index models y = σ * (⟨x, β⟩), which are low-dimensional nonlinear functions parameterized by feature vector β. We prove that transformers pretrained by gradient-based optimization can perform inference-time feature learning, i.e., extract information of the target features β solely from test prompts (despite β varying across different prompts), hence achieving an in-context statistical efficiency that surpasses any non-adaptive (fixed-basis) algorithms such as kernel methods. Moreover, we show that the inference-time sample complexity surpasses the Correlational Statistical Query (CSQ) lower bound, owing to nonlinear label transformations naturally induced by the Softmax self-attention mechanism.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-LearningTomoya Wakayama, Taiji SuzukiICML 2026 · 被引用 12 次
- From Shortcut to Induction Head: How Data Diversity Shapes Algorithm Selection in TransformersRyotaro Kawata, Yujin Song, Alberto Bietti, Naoki Nishikawa 等NeurIPS 2025 · 被引用 7 次
- Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax OptimalityRyotaro Kawata, Taiji SuzukiICLR 2026 · 被引用 3 次
- Dataset Distillation Efficiently Encodes Low-Dimensional Representations from Gradient-Based Learning of Non-Linear TasksYuri Kinoshita, Naoki Nishikawa, Taro ToyoizumiICML 2026
它引用的顶会 Paper36
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 被引用 883 次
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento 等ICML 2023 · 被引用 729 次
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong 等NeurIPS 2023 · 被引用 356 次
- Transformers learn to implement preconditioned gradient descent for in-context learningKwangjun Ahn, Xiang Cheng, Hadi Daneshmand, Suvrit SraNeurIPS 2023 · 被引用 324 次
相关 Paper
- Pretrained Transformer Efficiently Learns Low-Dimensional Target Functions In-ContextKazusato Oko, Yujin Song, Taiji Suzuki, Denny WuNeurIPS 2024 · 被引用 34 次
- How Many Pretraining Tasks Are Needed for In-Context Learning of Linear Regression?Jingfeng Wu, Difan Zou, Zixiang Chen, Vladimir Braverman 等ICLR 2024 · 被引用 94 次
- What learning algorithm is in-context learning? Investigations with linear modelsEkin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma 等ICLR 2023 · 被引用 85 次
- Efficient and Minimax Optimal In-context Nonparametric Regression with TransformersMichelle Ching, Ioana Popescu, Nico Smith, Tianyi Ma 等ICML 2026 · 被引用 6 次
- Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient DescentChenyang Zhang, Yuan CaoICML 2026 · 被引用 1 次
