Expect the Unexpected? Testing the Surprisal of Salient Entities
Jessica Lin, Amir Zeldes
Abstract
Previous work examining the Uniform Information Density (UID) hypothesis has shown that while information as measured by surprisal metrics is distributed more or less evenly across documents overall, local discrepancies can arise due to functional pressures corresponding to syntactic and discourse structural constraints. However, work thus far has largely disregarded the relative salience of discourse participants. We fill this gap by studying how overall salience of entities in discourse relates to surprisal using 70K manually annotated mentions across 16 genres of English and a novel minimal-pair prompting method. Our results show that globally salient entities exhibit significantly higher surprisal than non-salient ones, even controlling for position, length, and nesting confounds. Moreover, salient entities systematically reduce surprisal for surrounding content when used as prompts, enhancing document-level predictability. This effect varies by genre, appearing strongest in topic-coherent texts and weakest in conversational contexts. Our findings refine the UID competing pressures framework by identifying global entity salience as a mechanism shaping information distribution in discourse.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f75d24a-1c85-4500-b9e3-38d5d9ea2a79Builds on3
- Revisiting the Uniform Information Density HypothesisClara Meister, Tiago Pimentel, Patrick Haller, Lena A. Jäger et al.EMNLP 2021 · 4 citations
- Surprise! Uniform Information Density Isn't the Whole Story: Predicting Surprisal Contours in Long-form DiscourseEleftheria Tsipidi, Franz Nowak, Ryan Cotterell, Ethan Wilcox et al.EMNLP 2024 · 2 citations
- GDTB: Genre Diverse Data for English Shallow Discourse Parsing across Modalities, Text Types, and DomainsYang Janet Liu, Tatsuya Aoyama, Wesley Scivetti, Yilun Zhu et al.EMNLP 2024
Related papers
- A Cognitive Regularizer for Language ModelingJason Wei, Clara Meister, Ryan CotterellACL 2021
- Uniform Information Density and Syntactic Reduction: Revisiting that-Mentioning in English Complement ClausesHailin Hao, Elsi KaiserEMNLP 2025
- The Harmonic Structure of Information ContoursEleftheria Tsipidi, Samuel Kiegeland, Franz Nowak, Tianyang Xu et al.ACL 2025 · 6 citations
- Surprisal Minimisation over Goal-directed Alternatives Predicts Production Choice in DialogueThomas P. Utting, Mario Giulianelli, Arabella SinclairACL 2026
- Information Value: Measuring Utterance Predictability as Distance from Plausible AlternativesMario Giulianelli, Sarenne Wallbridge, Raquel FernándezEMNLP 2023 · 4 citations
