Identifying the Periodicity of Information in Natural Language
Yulin Ou, Yu Wang, Yang Xu, Hendrik Buschmeier
Abstract
Recent theoretical advancement of information density in natural language have raised the following question: To what degree does natural language exhibit periodicity pattern in its encoded information? We address this question by introducing a new method called AutoPeriod of Surprisal (APS). APS adopts a canonical periodicity detection algorithm and is able to identify any significant periods that exist in the surprisal sequence of a single document. By applying the algorithm to a set of corpora, we have obtained the following empirical results: Firstly, a considerable proportion of human language demonstrates a strong pattern of periodicity in information. Secondly, new periods that are outside the distributions of typical structural units in text (e.g., sentence boundaries, elementary discourse units, etc.) are found and further confirmed via harmonic regression modeling. We conclude that the periodicity of information in language is a joint outcome from both structured factors and other driving factors that take effect at longer distances. The advantages of our periodicity detection method and its potentials in LLM-generation detection are further discussed. Our code is available as public repositories. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e5139d8-ff0c-4c6f-aa06-97e989e1944aBuilds on6
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 508 citations
- Detecting Subtle Differences between Human and Model Languages Using Spectrum of Relative LikelihoodYang Xu, Yu Wang, Hao An, Zhichen Liu et al.EMNLP 2024 · 20 citations
- FACE: Evaluating Natural Language Generation with Fourier Analysis of Cross-EntropyZuhao Yang, Yingfang Yuan, Yang Xu, Shuo Zhan et al.NeurIPS 2023 · 11 citations
- The Harmonic Structure of Information ContoursEleftheria Tsipidi, Samuel Kiegeland, Franz Nowak, Tianyang Xu et al.ACL 2025 · 6 citations
- Revisiting the Uniform Information Density HypothesisClara Meister, Tiago Pimentel, Patrick Haller, Lena A. Jäger et al.EMNLP 2021 · 4 citations
Related papers
- Surprise! Uniform Information Density Isn't the Whole Story: Predicting Surprisal Contours in Long-form DiscourseEleftheria Tsipidi, Franz Nowak, Ryan Cotterell, Ethan Wilcox et al.EMNLP 2024 · 2 citations
- On the Proper Treatment of Units in Surprisal TheorySamuel Kiegeland, Vésteinn Snæbjarnarson, Tim Vieira, Ryan CotterellACL 2026
- AmortizedPeriod: Attention-based Amortized Inference for Periodicity IdentificationHang Yu, Cong Liao, Ruolan Liu, Jianguo Li et al.ICLR 2024 · 2 citations
- Expect the Unexpected? Testing the Surprisal of Salient EntitiesJessica Lin, Amir ZeldesACL 2026 · 2 citations
- Prior-based Noisy Text Data Filtering: Fast and Strong Alternative For PerplexityYeongbin Seo, Gayoung Kim, Jaehyung Kim, Jinyoung YeoICLR 2026
