The Harmonic Structure of Information Contours
Eleftheria Tsipidi, Samuel Kiegeland, Franz Nowak, Tianyang Xu, Ethan Wilcox, Alex Warstadt, Ryan Cotterell, Mario Giulianelli
摘要
The uniform information density (UID) hypothesis proposes that speakers aim to distribute information evenly throughout a text, balancing production effort and listener comprehension difficulty. However, language typically does not maintain a strictly uniform information rate; instead, it fluctuates around a global average. These fluctuations are often explained by factors such as syntactic constraints, stylistic choices, or audience design. In this work, we explore an alternative perspective: that these fluctuations may be influenced by an implicit linguistic pressure towards periodicity, where the information rate oscillates at regular intervals, potentially across multiple frequencies simultaneously. We apply harmonic regression and introduce a novel extension called time scaling to detect and test for such periodicity in information contours. Analyzing texts in English, Spanish, German, Dutch, Basque, and Brazilian Portuguese, we find consistent evidence of periodic patterns in information rate. Many dominant frequencies align with discourse structure, suggesting these oscillations reflect meaningful linguistic organization. Beyond highlighting the connection between information rate and discourse structure, our approach offers a general framework for uncovering structural pressures at various levels of linguistic granularity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Identifying the Periodicity of Information in Natural LanguageYulin Ou, Yu Wang, Yang Xu, Hendrik BuschmeierACL 2026 · 被引用 2 次
- Surprisal Minimisation over Goal-directed Alternatives Predicts Production Choice in DialogueThomas P. Utting, Mario Giulianelli, Arabella SinclairACL 2026
- On the Proper Treatment of Units in Surprisal TheorySamuel Kiegeland, Vésteinn Snæbjarnarson, Tim Vieira, Ryan CotterellACL 2026
它引用的顶会 Paper10
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 被引用 508 次
- Detecting Subtle Differences between Human and Model Languages Using Spectrum of Relative LikelihoodYang Xu, Yu Wang, Hao An, Zhichen Liu 等EMNLP 2024 · 被引用 20 次
- FACE: Evaluating Natural Language Generation with Fourier Analysis of Cross-EntropyZuhao Yang, Yingfang Yuan, Yang Xu, Shuo Zhan 等NeurIPS 2023 · 被引用 11 次
- Segment Any Text: A Universal Approach for Robust, Efficient and Adaptable Sentence SegmentationMarkus Frohmann, Igor Sterner, Ivan Vulic, Benjamin Minixhofer 等EMNLP 2024 · 被引用 10 次
- Where's the Point? Self-Supervised Multilingual Punctuation-Agnostic Sentence SegmentationBenjamin Minixhofer, Jonas Pfeiffer, Ivan VulicACL 2023 · 被引用 8 次
相关 Paper
- Surprise! Uniform Information Density Isn't the Whole Story: Predicting Surprisal Contours in Long-form DiscourseEleftheria Tsipidi, Franz Nowak, Ryan Cotterell, Ethan Wilcox 等EMNLP 2024 · 被引用 2 次
- Revisiting the Uniform Information Density HypothesisClara Meister, Tiago Pimentel, Patrick Haller, Lena A. Jäger 等EMNLP 2021 · 被引用 4 次
- A Cognitive Regularizer for Language ModelingJason Wei, Clara Meister, Ryan CotterellACL 2021
- Uniform Information Density and Syntactic Reduction: Revisiting that-Mentioning in English Complement ClausesHailin Hao, Elsi KaiserEMNLP 2025
- Expect the Unexpected? Testing the Surprisal of Salient EntitiesJessica Lin, Amir ZeldesACL 2026 · 被引用 2 次
