Bridging Information-Theoretic and Geometric Compression in Language Models
Emily Cheng, Corentin Kervadec, Marco Baroni
摘要
For a language model (LM) to faithfully model human language, it must compress vast, potentially infinite information into relatively few dimensions. We propose analyzing compression in (pre-trained) LMs from two points of view: geometric and information-theoretic. We demonstrate that the two views are highly correlated, such that the intrinsic geometric dimension of linguistic data predicts their coding length under the LM. We then show that, in turn, high compression of a linguistic dataset predicts rapid adaptation to that dataset, confirming that being able to compress linguistic information is an important part of successful LM performance. As a practical byproduct of our analysis, we evaluate a battery of intrinsic dimension estimators for the first time on linguistic data, showing that only some encapsulate the relationship between information-theoretic compression, geometric compression, and ease-of-adaptation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- The geometry of hidden representations of large transformer modelsLucrezia Valeriani, Diego Doimo, Francesca Cuturello, Alessandro Laio 等NeurIPS 2023 · 被引用 148 次
- The Representation Landscape of Few-Shot Learning and Fine-Tuning in Large Language ModelsDiego Doimo, Alessandro Serra, Alessio Ansuini, Alberto CazzanigaNeurIPS 2024 · 被引用 21 次
- Intrinsic Entropy of Context Length Scaling in LLMsJingzhe Shi, Qinwei Ma, Hongyi Liu, Hang Zhao 等ICLR 2026 · 被引用 17 次
- Geometry of Decision Making in Language ModelsAbhinav Joshi, Divyanshu Bhatt, Ashutosh ModiNeurIPS 2025 · 被引用 12 次
- Dimension Importance Estimation for Dense Information RetrievalGuglielmo Faggioli, Nicola Ferro, Raffaele Perego, Nicola TonellottoSIGIR 2024 · 被引用 7 次
它引用的顶会 Paper7
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- The Intrinsic Dimension of Images and Its Impact on LearningPhillip Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum 等ICLR 2021 · 被引用 381 次
- The geometry of hidden representations of large transformer modelsLucrezia Valeriani, Diego Doimo, Francesca Cuturello, Alessandro Laio 等NeurIPS 2023 · 被引用 148 次
- Emergence of Separable Manifolds in Deep Language RepresentationsJonathan Mamou, Hang Le, Miguel Del Rio, Cory Stephenson 等ICML 2020 · 被引用 52 次
- Information-Theoretic Probing with Minimum Description LengthElena Voita, Ivan TitovEMNLP 2020 · 被引用 34 次
相关 Paper
- Geometric Signatures of Compositionality Across a Language Model's LifetimeJin Hwa Lee, Thomas Jiralerspong, Lei Yu, Yoshua Bengio 等ACL 2025
- Emergence of a High-Dimensional Abstraction Phase in Language TransformersEmily Cheng, Diego Doimo, Corentin Kervadec, Iuri Macocco 等ICLR 2025 · 被引用 1 次
- Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-TuningArmen Aghajanyan, Sonal Gupta, Luke ZettlemoyerACL 2021
- Learning is Forgetting; LLM Training As Lossy CompressionHenry Conklin, Tom Hosking, Yi Chern Tan, Jonathan D. Cohen 等ICLR 2026 · 被引用 6 次
- Abstraction Induces the Brain Alignment of Language and Speech ModelsEmily Cheng, Aditya Vaidya, Richard AntonelloICML 2026
