Is attention required for ICL? Exploring the Relationship Between Model Architecture and In-Context Learning Ability
Ivan Lee, Nan Jiang, Taylor Berg-Kirkpatrick
Abstract
What is the relationship between model architecture and the ability to perform in-context learning? In this empirical study, we take the first steps toward answering this question. We evaluate thirteen model architectures capable of causal language modeling across a suite of synthetic in-context learning tasks. These selected architectures represent a broad range of paradigms, including recurrent and convolution-based neural networks, transformers, state space model inspired, and other emerging attention alternatives. We discover that all the considered architectures can perform in-context learning under a wider range of conditions than previously documented. Additionally, we observe stark differences in statistical efficiency and consistency by varying the number of in-context examples and task difficulty. We also measure each architecture's predisposition towards in-context learning when presented with the option to memorize rather than leverage in-context examples. Finally, and somewhat surprisingly, we find that several attention alternatives are sometimes competitive with or better in-context learners than transformers. However, no single architecture demonstrates consistency across all tasks, with performance either plateauing or declining when confronted with a significantly larger number of in-context examples than those encountered during gradient-based training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 05bcadc3-47e2-4228-96cf-5cafafb44631Cited by top-tier papers5
- Linking In-context Learning in Transformers to Human Episodic MemoryJi-An Li, Corey Yishan Zhou, Marcus K. Benna, Marcelo G. MattarNeurIPS 2024 · 20 citations
- Weight-Space Linear Recurrent Neural NetworksRoussel Desmond Nzoyem, Nawid Keshtmand, Enrique Crespo-Fernandez, Idriss Tsayem et al.ICLR 2026 · 6 citations
- Trained Mamba Emulates Online Gradient Descent in In-Context Linear RegressionJiarui Jiang, Wei Huang, Miao Zhang, Taiji Suzuki et al.NeurIPS 2025 · 2 citations
- Training Dynamics of In-Context Learning in Linear AttentionYedi Zhang, Aaditya K. Singh, Peter E. Latham, Andrew M. SaxeICML 2025
- Train Once, Reuse Everywhere: Generalizable Implicit ICL by Routing AttentionJiaqian Li, Yanshu Li, Ligong Han, Ruixiang Tang et al.ICML 2026
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
Related papers
- CausalLM is not optimal for in-context learningNan Ding, Tomer Levinboim, Jialin Wu, Sebastian Goodman et al.ICLR 2024 · 35 citations
- BERTs are Generative In-Context LearnersDavid SamuelNeurIPS 2024 · 17 citations
- MLPs Learn In-Context on Regression and Classification TasksWilliam Lingxiao Tong, Cengiz PehlevanICLR 2025
- Differential learning kinetics govern the transition from memorization to generalization during in-context learningAlex Nguyen, Gautam ReddyICLR 2025
- Data Distributional Properties Drive Emergent In-Context Learning in TransformersStephanie C. Y. Chan, Adam Santoro, Andrew K. Lampinen, Jane X. Wang et al.NeurIPS 2022 · 407 citations
