Efficient and Minimax Optimal In-context Nonparametric Regression with Transformers
Michelle Ching, Ioana Popescu, Nico Smith, Tianyi Ma, William Underwood, Richard Samworth
摘要
We study in-context learning for nonparametric regression with -Hölder smooth regression functions, for some . We prove that, with in-context examples and -dimensional regression covariates, a pretrained transformer with parameters and pretraining sequences can achieve the minimax optimal rate of convergence in mean squared error. Our result requires substantially fewer transformer parameters and pretraining sequences than previous results in the literature. This is achieved by showing that transformers are able to approximate local polynomial estimators efficiently by implementing a kernel-weighted polynomial basis and then running gradient descent.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 被引用 883 次
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm SelectionYu Bai, Fan Chen, Huan Wang, Caiming Xiong 等NeurIPS 2023 · 被引用 356 次
- Transformers learn to implement preconditioned gradient descent for in-context learningKwangjun Ahn, Xiang Cheng, Hadi Daneshmand, Suvrit SraNeurIPS 2023 · 被引用 324 次
相关 Paper
- Transformers are Minimax Optimal Nonparametric In-Context LearnersJuno Kim, Tai Nakamaki, Taiji SuzukiNeurIPS 2024 · 被引用 42 次
- Understanding In-Context Learning on Structured Manifolds: Bridging Attention to Kernel MethodsZhaiming Shen, Alexander Hsu, Rongjie Lai, Wenjing LiaoICLR 2026 · 被引用 13 次
- What learning algorithm is in-context learning? Investigations with linear modelsEkin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma 等ICLR 2023 · 被引用 85 次
- Pretrained Transformer Efficiently Learns Low-Dimensional Target Functions In-ContextKazusato Oko, Yujin Song, Taiji Suzuki, Denny WuNeurIPS 2024 · 被引用 34 次
- Approximation Bounds for Transformer Networks with Application to RegressionYuling Jiao, Yanming Lai, Defeng Sun, Yang Wang 等ICML 2026 · 被引用 6 次
