Language Modeling Is Compression
Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, Marcus Hutter, Joel Veness
摘要
It has long been established that predictive models can be transformed into lossless compressors and vice versa. Incidentally, in recent years, the machine learning community has focused on training increasingly large and powerful self-supervised (language) models. Since these large language models exhibit impressive predictive capabilities, they are well-positioned to be strong compressors. In this work, we advocate for viewing the prediction problem through the lens of compression and evaluate the compression capabilities of large (foundation) models. We show that large language models are powerful general-purpose predictors and that the compression viewpoint provides novel insights into scaling laws, tokenization, and in-context learning. For example, Chinchilla 70B, while trained primarily on text, compresses ImageNet patches to 43.4% and LibriSpeech samples to 16.4% of their raw size, beating domain-specific compressors like PNG (58.5%) or FLAC (30.3%), respectively. Finally, we show that the prediction-compression equivalence allows us to use any compressor (like gzip) to build a conditional generative model. * Equal contribution. 1 Google DeepMind.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper89
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang 等NeurIPS 2025 · 被引用 949 次
- Large Language Models Are Zero-Shot Time Series ForecastersNate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon WilsonNeurIPS 2023 · 被引用 898 次
- In-context Autoencoder for Context Compression in a Large Language ModelTao Ge, Jing Hu, Lei Wang, Xun Wang 等ICLR 2024 · 被引用 158 次
- Rethinking LLM Memorization through the Lens of Adversarial CompressionAvi Schwarzschild, Zhili Feng, Pratyush Maini, Zachary C. Lipton 等NeurIPS 2024 · 被引用 120 次
- Fine-Tuned Language Models Generate Stable Inorganic Materials as TextNate Gruver, Anuroop Sriram, Andrea Madotto, Andrew Gordon Wilson 等ICLR 2024 · 被引用 120 次
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Compression of Generative Pre-trained Language Models via QuantizationChaofan Tao, Lu Hou, Wei Zhang, Lifeng Shang 等ACL 2022 · 被引用 119 次
- TRACE: A Fast Transformer-based General-Purpose Lossless CompressorYu Mao, Yufei Cui, Tei-Wei Kuo, Chun Jason XueWWW 2022 · 被引用 60 次
相关 Paper
- Compression via Pre-trained Transformers: A Study on Byte-Level Multimodal DataDavid Heurtel-Depeiges, Anian Ruoss, Joel Veness, Tim GeneweinICML 2025
- Large Language Models for Lossless Image Compression: Next-Pixel Prediction in Language Space is All You NeedKecheng Chen, Pingping Zhang, Hui Liu, Jie Liu 等NeurIPS 2025 · 被引用 14 次
- Unlocking Tokens as Data Points for Generalization Bounds on Larger Language ModelsSanae Lotfi, Yilun Kuang, Marc Finzi, Brandon Amos 等NeurIPS 2024 · 被引用 29 次
- Non-Vacuous Generalization Bounds for Large Language ModelsSanae Lotfi, Marc Anton Finzi, Yilun Kuang, Tim G. J. Rudner 等ICML 2024 · 被引用 49 次
- Learning is Forgetting; LLM Training As Lossy CompressionHenry Conklin, Tom Hosking, Yi Chern Tan, Jonathan D. Cohen 等ICLR 2026 · 被引用 6 次
