A Tale of Two Structures: Do LLMs Capture the Fractal Complexity of Language?
Ibrahim Alabdulmohsin, Andreas Peter Steiner
Abstract
Language exhibits a fractal structure in its information-theoretic complexity (i.e. bits per token), with self-similarity across scales and longrange dependence (LRD). In this work, we investigate whether large language models (LLMs) can replicate such fractal characteristics and identify conditions-such as temperature setting and prompting method-under which they may fail. Moreover, we find that the fractal parameters observed in natural language are contained within a narrow range, whereas those of LLMs' output vary widely, suggesting that fractal parameters might prove helpful in detecting a non-trivial portion of LLM-generated texts. Notably, these findings, and many others reported in this work, are robust to the choice of the architecture; e.g. Gemini 1.0 Pro, Mistral-7B and Gemma-2B. We also release a dataset comprising over 240,000 articles generated by various LLMs (both pretrained and instruction-tuned) with different decoding temperatures and prompting methods, along with their corresponding human-generated texts. We hope that this work highlights the complex interplay between fractal properties, prompting, and statistical mimicry in LLMs, offering insights for generating, evaluating and detecting synthetic texts. Ask the model to generate an outline before generating the article, using a short prefix. short keywords (kw) A few, unordered keywords. keywords (kw+) Many unordered keywords.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c12663b8-c5e7-4e99-bfe2-4686a1931ac0Cited by top-tier papers1
Ask how each one uses itBuilds on15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning et al.ICML 2023 · 988 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated TextAbhimanyu Hans, Avi Schwarzschild, Valeriia Cherepanova, Hamid Kazemi et al.ICML 2024 · 262 citations
- DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated TextXianjun Yang, Wei Cheng, Yue Wu, Linda Ruth Petzold et al.ICLR 2024 · 173 citations
Related papers
- Fractal Patterns May Illuminate the Success of Next-Token PredictionIbrahim M. Alabdulmohsin, Vinh Q. Tran, Mostafa DehghaniNeurIPS 2024 · 8 citations
- Correlation Dimension of Autoregressive Large Language ModelsXin Du, Kumiko Tanaka-IshiiNeurIPS 2025 · 2 citations
- Understanding and Mitigating Language Confusion in LLMsKelly Marchisio, Wei-Yin Ko, Alexandre Berard, Théo Dehaze et al.EMNLP 2024 · 10 citations
- Measuring Intent Comprehension in LLMsNadav Kunievsky, James EvansICML 2026 · 1 citation
- D-Models and E-Models: Diversity-Stability Trade-offs in the Sampling Behavior of Large Language ModelsJia Gu, Liang Pang, Huawei Shen, Xueqi ChengWWW 2026
