Correlation Dimension of Autoregressive Large Language Models
Xin Du, Kumiko Tanaka-Ishii
摘要
Large language models (LLMs) have achieved remarkable progress in natural language generation, yet they continue to display puzzling behaviors -- such as repetition and incoherence -- even when exhibiting low perplexity. This highlights a key limitation of conventional evaluation metrics, which emphasize local prediction accuracy while overlooking long-range structural complexity. We introduce correlation dimension, a fractal-geometric measure of self-similarity, to quantify the epistemological complexity of text as perceived by a language model. This measure captures the hierarchical recurrence structure of language, bridging local and global properties in a unified framework. Through extensive experiments, we show that correlation dimension (1) reveals three distinct phases during pretraining, (2) reflects context-dependent complexity, (3) indicates a model's tendency toward hallucination, and (4) reliably detects multiple forms of degeneration in generated text. The method is computationally efficient, robust to model quantization (down to 4-bit precision), broadly applicable across autoregressive architectures (e.g., Transformer and Mamba), and provides fresh insight into the generative dynamics of LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Context parroting: A simple but tough-to-beat baseline for foundation models in scientific machine learningYuanzhao Zhang, William GilpinICLR 2026 · 被引用 16 次
- A universal compression theory for lottery ticket hypothesis and neural scaling lawsHong-Yi Wang, Di Luo, Tomaso Poggio, Isaac L. Chuang 等ICLR 2026 · 被引用 2 次
- Escaping Mode Collapse in LLM Generation via Geometric RegulationXin Du, Kumiko Tanaka-IshiiICML 2026
它引用的顶会 Paper9
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence FrontiersKrishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun 等NeurIPS 2021 · 被引用 606 次
相关 Paper
- Repeated Sequences Reveal Gaps between Large Language Models and Natural LanguageKumiko Tanaka-IshiiACL 2026 · 被引用 1 次
- Fractal Patterns May Illuminate the Success of Next-Token PredictionIbrahim M. Alabdulmohsin, Vinh Q. Tran, Mostafa DehghaniNeurIPS 2024 · 被引用 8 次
- Characterizing Truthfulness in Large Language Model Generations with Local Intrinsic DimensionFan Yin, Jayanth Srinivasa, Kai-Wei ChangICML 2024 · 被引用 43 次
- A Tale of Two Structures: Do LLMs Capture the Fractal Complexity of Language?Ibrahim Alabdulmohsin, Andreas Peter SteinerICML 2025
- Geometric Signatures of Compositionality Across a Language Model's LifetimeJin Hwa Lee, Thomas Jiralerspong, Lei Yu, Yoshua Bengio 等ACL 2025
