Is attention required for ICL? Exploring the Relationship Between Model Architecture and In-Context Learning Ability
Ivan Lee, Nan Jiang, Taylor Berg-Kirkpatrick
摘要
What is the relationship between model architecture and the ability to perform in-context learning? In this empirical study, we take the first steps toward answering this question. We evaluate thirteen model architectures capable of causal language modeling across a suite of synthetic in-context learning tasks. These selected architectures represent a broad range of paradigms, including recurrent and convolution-based neural networks, transformers, state space model inspired, and other emerging attention alternatives. We discover that all the considered architectures can perform in-context learning under a wider range of conditions than previously documented. Additionally, we observe stark differences in statistical efficiency and consistency by varying the number of in-context examples and task difficulty. We also measure each architecture's predisposition towards in-context learning when presented with the option to memorize rather than leverage in-context examples. Finally, and somewhat surprisingly, we find that several attention alternatives are sometimes competitive with or better in-context learners than transformers. However, no single architecture demonstrates consistency across all tasks, with performance either plateauing or declining when confronted with a significantly larger number of in-context examples than those encountered during gradient-based training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Linking In-context Learning in Transformers to Human Episodic MemoryJi-An Li, Corey Yishan Zhou, Marcus K. Benna, Marcelo G. MattarNeurIPS 2024 · 被引用 20 次
- Weight-Space Linear Recurrent Neural NetworksRoussel Desmond Nzoyem, Nawid Keshtmand, Enrique Crespo-Fernandez, Idriss Tsayem 等ICLR 2026 · 被引用 6 次
- Trained Mamba Emulates Online Gradient Descent in In-Context Linear RegressionJiarui Jiang, Wei Huang, Miao Zhang, Taiji Suzuki 等NeurIPS 2025 · 被引用 2 次
- Training Dynamics of In-Context Learning in Linear AttentionYedi Zhang, Aaditya K. Singh, Peter E. Latham, Andrew M. SaxeICML 2025
- Train Once, Reuse Everywhere: Generalizable Implicit ICL by Routing AttentionJiaqian Li, Yanshu Li, Ligong Han, Ruixiang Tang 等ICML 2026
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein 等ICML 2021 · 被引用 1,843 次
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 被引用 1,168 次
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 被引用 1,030 次
相关 Paper
- CausalLM is not optimal for in-context learningNan Ding, Tomer Levinboim, Jialin Wu, Sebastian Goodman 等ICLR 2024 · 被引用 35 次
- BERTs are Generative In-Context LearnersDavid SamuelNeurIPS 2024 · 被引用 17 次
- MLPs Learn In-Context on Regression and Classification TasksWilliam Lingxiao Tong, Cengiz PehlevanICLR 2025
- Differential learning kinetics govern the transition from memorization to generalization during in-context learningAlex Nguyen, Gautam ReddyICLR 2025
- Data Distributional Properties Drive Emergent In-Context Learning in TransformersStephanie C. Y. Chan, Adam Santoro, Andrew K. Lampinen, Jane X. Wang 等NeurIPS 2022 · 被引用 407 次
