Information Integration in Large Language Models is Gated by Linguistic Structural Markers
Wei Liu, Nai Ding
摘要
Language comprehension relies on integrating information across both local words and broader context. We propose a method to quantify the information integration window of large language models (LLMs) and examine how sentence and clause boundaries constrain this window. Specifically, LLMs are required to predict a target word based on either a local window (local prediction) or the full context (global prediction), and we use Jensen-Shannon (JS) divergence to measure the information loss from relying solely on the local window, termed the local-prediction deficit. Results show that integration windows of both humans and LLMs are strongly modulated by sentence boundaries, and predictions primarily rely on words within the same sentence or clause: The localprediction deficit follows a power-law decay as the window length increases and drops sharply at the sentence boundary. This boundary effect is primarily attributed to linguistic structural markers, e.g., punctuation, rather than implicit syntactic or semantic cues. Together, these results indicate that LLMs rely on explicit structural cues to guide their information integration strategy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionJingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo 等ACL 2025 · 被引用 334 次
- Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERTZhiyong Wu, Yun Chen, Ben Kao, Qun LiuACL 2020 · 被引用 158 次
- LooGLE: Can Long-Context Language Models Understand Long Contexts?Jiaqi Li, Mengmeng Wang, Zilong Zheng, Muhan ZhangACL 2024 · 被引用 32 次
- Unsupervised Parsing via Constituency TestsSteven Cao, Nikita Kitaev, Dan KleinEMNLP 2020 · 被引用 25 次
相关 Paper
- Large language models transition from integrating across position-yoked, exponential windows to structure-yoked, power-law windowsDavid Skrill, Samuel Norman-HaignereNeurIPS 2023 · 被引用 7 次
- Prompting is not a substitute for probability measurements in large language modelsJennifer Hu, Roger LevyEMNLP 2023 · 被引用 31 次
- Comparing human and language models sentence processing difficulties on complex structuresSamuel Joseph Amouyal, Aya Meltzer-Asscher, Jonathan BerantACL 2026 · 被引用 1 次
- From Tokens to Thoughts: How LLMs and Humans Trade Compression for MeaningChen Shani, Liron Soffer, Dan Jurafsky, Yann LeCun 等ICLR 2026 · 被引用 38 次
- Clozing the Gap: Exploring Why Language Model Surprisal Outperforms Cloze SurprisalSathvik Nair, Byung-Doh OhACL 2026 · 被引用 1 次
