In-Context Learning and Occam's Razor
Eric Elmoznino, Tom Marty, Tejas Kasetty, Léo Gagnon, Sarthak Mittal, Mahan Fathi, Dhanya Sridhar, Guillaume Lajoie
摘要
A central goal of machine learning is generalization. While the No Free Lunch Theorem states that we cannot obtain theoretical guarantees for generalization without further assumptions, in practice we observe that simple models which explain the training data generalize best-a principle called Occam's razor. Despite the need for simple models, most current approaches in machine learning only minimize the training error, and at best indirectly promote simplicity through regularization or architecture design. Here, we draw a connection between Occam's razor and in-context learning-an emergent ability of certain sequence models like Transformers to learn at inference time from past observations in a sequence. In particular, we show that the next-token prediction loss used to train in-context learners is directly equivalent to a data compression technique called prequential coding, and that minimizing this loss amounts to jointly minimizing both the training error and the complexity of the model that was implicitly learned from context. Our theory and the empirical experiments we use to support it not only provide a normative account of in-context learning, but also elucidate the shortcomings of current in-context learning methods, suggesting ways in which they can be improved. We make our code available at https://github.com/ 3rdCore/PrequentialCode .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Multiple Token Divergence: Measuring and Steering In-Context Computation DensityVincent Herrmann, Eric Alcaide, Michael Wand, Jürgen SchmidhuberICLR 2026 · 被引用 1 次
- Technical Debt in In-Context Learning: Diminishing Efficiency in Long ContextTaejong Joo, Diego KlabjanNeurIPS 2025
- Measuring In-Context Computation Complexity via Hidden State PredictionVincent Herrmann, Róbert Csordás, Jürgen SchmidhuberICML 2025
它引用的顶会 Paper23
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 被引用 1,030 次
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 被引用 883 次
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento 等ICML 2023 · 被引用 729 次
- Data Distributional Properties Drive Emergent In-Context Learning in TransformersStephanie C. Y. Chan, Adam Santoro, Andrew K. Lampinen, Jane X. Wang 等NeurIPS 2022 · 被引用 407 次
相关 Paper
- In-Context Learning Strategies Emerge RationallyDaniel Wurgaft, Ekdeep Singh Lubana, Core Francisco Park, Hidenori Tanaka 等NeurIPS 2025 · 被引用 19 次
- Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for TransformersPeter Shaw, James Cohan, Jacob Eisenstein, Kristina ToutanovaICLR 2026 · 被引用 7 次
- What learning algorithm is in-context learning? Investigations with linear modelsEkin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma 等ICLR 2023 · 被引用 85 次
- PAC-Bayes Compression Bounds So Tight That They Can Explain GeneralizationSanae Lotfi, Marc Finzi, Sanyam Kapoor, Andres Potapczynski 等NeurIPS 2022 · 被引用 98 次
- Learning is Forgetting; LLM Training As Lossy CompressionHenry Conklin, Tom Hosking, Yi Chern Tan, Jonathan D. Cohen 等ICLR 2026 · 被引用 6 次
