Non-Vacuous Generalization Bounds for Large Language Models
Sanae Lotfi, Marc Anton Finzi, Yilun Kuang, Tim G. J. Rudner, Micah Goldblum, Andrew Gordon Wilson
摘要
Modern language models can contain billions of parameters, raising the question of whether they can generalize beyond the training data or simply parrot their training corpora. We provide the first non-vacuous generalization bounds for pretrained large language models (LLMs), indicating that language models are capable of discovering regularities that generalize to unseen data. In particular, we derive a compression bound that is valid for the unbounded log-likelihood loss using prediction smoothing, and we extend the bound to handle subsampling, accelerating bound computation by orders of magnitude on massive datasets. To achieve the extreme level of compression required for non-vacuous bounds, we devise SubLoRA, a simple low-dimensional nonlinear parameterization that leads to non-vacuous generalization bounds for models with nearly a billion parameters. Finally, we use our bounds to understand LLM generalization and find that larger models have better generalization bounds and are more compressible than smaller models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Unlocking Tokens as Data Points for Generalization Bounds on Larger Language ModelsSanae Lotfi, Yilun Kuang, Marc Finzi, Brandon Amos 等NeurIPS 2024 · 被引用 29 次
- The Coverage Principle: How Pre-Training Enables Post-TrainingFan Chen, Audrey Huang, Noah Golowich, Sadhika Malladi 等ICLR 2026 · 被引用 28 次
- Predicting the Performance of Black-box Language Models with Follow-up QueriesDylan Sam, Marc Finzi, Zico KolterNeurIPS 2025 · 被引用 10 次
- Can DPO Learn Diverse Human Values? A Theoretical Scaling LawShawn Im, Sharon LiNeurIPS 2025 · 被引用 8 次
- Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for TransformersPeter Shaw, James Cohan, Jacob Eisenstein, Kristina ToutanovaICLR 2026 · 被引用 7 次
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- Language Modeling Is CompressionGrégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt 等ICLR 2024 · 被引用 243 次
- Quantifying Memorization Across Neural Language ModelsNicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee 等ICLR 2023 · 被引用 158 次
相关 Paper
- Learning is Forgetting; LLM Training As Lossy CompressionHenry Conklin, Tom Hosking, Yi Chern Tan, Jonathan D. Cohen 等ICLR 2026 · 被引用 6 次
- Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-TuningArmen Aghajanyan, Sonal Gupta, Luke ZettlemoyerACL 2021
- Compute-Optimal LLMs Provably Generalize Better with ScaleMarc Anton Finzi, Sanyam Kapoor, Diego Granziol, Anming Gu 等ICLR 2025
- Radio: Rate-Distortion Optimization for Large Language Model CompressionSean I. YoungICML 2025
- SoLA: Leveraging Soft Activation Sparsity and Low-Rank Decomposition for Large Language Model CompressionXinhao Huang, You-Liang Huang, Zeyi WenAAAI 2025 · 被引用 14 次
