Surprise! Uniform Information Density Isn't the Whole Story: Predicting Surprisal Contours in Long-form Discourse
Eleftheria Tsipidi, Franz Nowak, Ryan Cotterell, Ethan Wilcox, Mario Giulianelli, Alex Warstadt
摘要
The Uniform Information Density (UID) hypothesis posits that speakers tend to distribute information evenly across linguistic units to achieve efficient communication. Of course, information rate in texts and discourses is not perfectly uniform. While these fluctuations can be viewed as theoretically uninteresting noise on top of a uniform target, another explanation is that UID is not the only functional pressure regulating information content in a language. Speakers may also seek to maintain interest, adhere to writing conventions, and build compelling arguments. In this paper, we propose one such functional pressure; namely that speakers modulate information rate based on location within a hierarchically-structured model of discourse. We term this the Structured Context Hypothesis and test it by predicting the surprisal contours of naturally occurring discourses extracted from large language models using predictors derived from discourse structure. We find that hierarchical predictors are significant predictors of a discourse's information contour and that deeply nested hierarchical predictors are more predictive than shallow ones. This work takes an initial step beyond UID to propose testable hypotheses for why the information rate fluctuates in predictable ways.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Anything Goes? A Crosslinguistic Study of (Im)possible Language Learning in LMsXiulin Yang, Tatsuya Aoyama, Yuekun Yao, Ethan WilcoxACL 2025 · 被引用 9 次
- The Harmonic Structure of Information ContoursEleftheria Tsipidi, Samuel Kiegeland, Franz Nowak, Tianyang Xu 等ACL 2025 · 被引用 6 次
- Expect the Unexpected? Testing the Surprisal of Salient EntitiesJessica Lin, Amir ZeldesACL 2026 · 被引用 2 次
- Identifying the Periodicity of Information in Natural LanguageYulin Ou, Yu Wang, Yang Xu, Hendrik BuschmeierACL 2026 · 被引用 2 次
- From Input Perception to Predictive Insight: Modeling Model Blind Spots Before They Become ErrorsMaggie Mi, Aline Villavicencio, Nafise Sadat MoosaviEMNLP 2025
它引用的顶会 Paper3
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 被引用 508 次
- Revisiting the Uniform Information Density HypothesisClara Meister, Tiago Pimentel, Patrick Haller, Lena A. Jäger 等EMNLP 2021 · 被引用 4 次
- An Exploration of Left-Corner TransformationsAndreas Opedal, Eleftheria Tsipidi, Tiago Pimentel, Ryan Cotterell 等EMNLP 2023
相关 Paper
- A Cognitive Regularizer for Language ModelingJason Wei, Clara Meister, Ryan CotterellACL 2021
- Uniform Information Density and Syntactic Reduction: Revisiting that-Mentioning in English Complement ClausesHailin Hao, Elsi KaiserEMNLP 2025
- Surprisal Minimisation over Goal-directed Alternatives Predicts Production Choice in DialogueThomas P. Utting, Mario Giulianelli, Arabella SinclairACL 2026
- Discourse Context Predictability Effects in Hindi Word OrderSidharth Ranjan, Marten van Schijndel, Sumeet Agarwal, Rajakrishnan RajkumarEMNLP 2022
- A Top-down Neural Architecture towards Text-level Parsing of Discourse Rhetorical StructureLongyin Zhang, Yuqing Xing, Fang Kong, Peifeng Li 等ACL 2020 · 被引用 39 次
