Identifying the Periodicity of Information in Natural Language
Yulin Ou, Yu Wang, Yang Xu, Hendrik Buschmeier
摘要
Recent theoretical advancement of information density in natural language have raised the following question: To what degree does natural language exhibit periodicity pattern in its encoded information? We address this question by introducing a new method called AutoPeriod of Surprisal (APS). APS adopts a canonical periodicity detection algorithm and is able to identify any significant periods that exist in the surprisal sequence of a single document. By applying the algorithm to a set of corpora, we have obtained the following empirical results: Firstly, a considerable proportion of human language demonstrates a strong pattern of periodicity in information. Secondly, new periods that are outside the distributions of typical structural units in text (e.g., sentence boundaries, elementary discourse units, etc.) are found and further confirmed via harmonic regression modeling. We conclude that the periodicity of information in language is a joint outcome from both structured factors and other driving factors that take effect at longer distances. The advantages of our periodicity detection method and its potentials in LLM-generation detection are further discussed. Our code is available as public repositories. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 被引用 508 次
- Detecting Subtle Differences between Human and Model Languages Using Spectrum of Relative LikelihoodYang Xu, Yu Wang, Hao An, Zhichen Liu 等EMNLP 2024 · 被引用 20 次
- FACE: Evaluating Natural Language Generation with Fourier Analysis of Cross-EntropyZuhao Yang, Yingfang Yuan, Yang Xu, Shuo Zhan 等NeurIPS 2023 · 被引用 11 次
- The Harmonic Structure of Information ContoursEleftheria Tsipidi, Samuel Kiegeland, Franz Nowak, Tianyang Xu 等ACL 2025 · 被引用 6 次
- Revisiting the Uniform Information Density HypothesisClara Meister, Tiago Pimentel, Patrick Haller, Lena A. Jäger 等EMNLP 2021 · 被引用 4 次
相关 Paper
- Surprise! Uniform Information Density Isn't the Whole Story: Predicting Surprisal Contours in Long-form DiscourseEleftheria Tsipidi, Franz Nowak, Ryan Cotterell, Ethan Wilcox 等EMNLP 2024 · 被引用 2 次
- On the Proper Treatment of Units in Surprisal TheorySamuel Kiegeland, Vésteinn Snæbjarnarson, Tim Vieira, Ryan CotterellACL 2026
- AmortizedPeriod: Attention-based Amortized Inference for Periodicity IdentificationHang Yu, Cong Liao, Ruolan Liu, Jianguo Li 等ICLR 2024 · 被引用 2 次
- Expect the Unexpected? Testing the Surprisal of Salient EntitiesJessica Lin, Amir ZeldesACL 2026 · 被引用 2 次
- Prior-based Noisy Text Data Filtering: Fast and Strong Alternative For PerplexityYeongbin Seo, Gayoung Kim, Jaehyung Kim, Jinyoung YeoICLR 2026
