Learning is Forgetting; LLM Training As Lossy Compression
Henry Conklin, Tom Hosking, Yi Chern Tan, Jonathan D. Cohen, Sarah-Jane Leslie, Thomas L. Griffiths, Max Bartolo, Seraphina Goldfarb-Tarrant
摘要
Despite the increasing prevalence of large language models (LLMs), we still have a limited understanding of how their representational spaces are structured. This limits our ability to interpret how and what they learn or relate them to learning in humans. We argue LLMs are best seen as an instance of lossy compression, where over training they learn by retaining only information in their training data relevant to their objective(s). We show pre-training results in models that are optimally compressed for next-sequence prediction, approaching the Information Bottleneck bound on compression. Across an array of open weights models, each compresses differently, likely due to differences in the data and training recipes used. However even across different families of LLMs the optimality of a model's compression, and the information present in it, can predict downstream performance on across a wide array of benchmarks, letting us directly link representational structure to actionable insights about model performance. In the general case the work presented here offers a unified Information-Theoretic framing for how these models learn that is deployable at scale.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action ModelYixu Feng, Zinan Zhao, Yanxiang Ma, Chenghao Xia 等ICML 2026 · 被引用 6 次
- The Geometry of Representational Failures in Vision Language ModelsDaniele Savietto, Declan Campbell, André Panisson, Marco Nurisso 等ICML 2026 · 被引用 5 次
它引用的顶会 Paper12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
相关 Paper
- Learning to Compress: Unlocking the Potential of Large Language Models for Text RepresentationYeqin Zhang, Yizheng Zhao, Chen Hu, Binxing Jiao 等AAAI 2026 · 被引用 2 次
- Non-Vacuous Generalization Bounds for Large Language ModelsSanae Lotfi, Marc Anton Finzi, Yilun Kuang, Tim G. J. Rudner 等ICML 2024 · 被引用 49 次
- Unlocking Tokens as Data Points for Generalization Bounds on Larger Language ModelsSanae Lotfi, Yilun Kuang, Marc Finzi, Brandon Amos 等NeurIPS 2024 · 被引用 29 次
- From Tokens to Thoughts: How LLMs and Humans Trade Compression for MeaningChen Shani, Liron Soffer, Dan Jurafsky, Yann LeCun 等ICLR 2026 · 被引用 38 次
- Representation Learning with Conditional Information Flow MaximizationDou Hu, Lingwei Wei, Wei Zhou, Songlin HuACL 2024
