Dynamics of Spontaneous Topic Changes in Next Token Prediction with Self-Attention
Mumin Jia, Jairo Diaz Rodriguez
Abstract
Human cognition is punctuated by abrupt, spontaneous shifts between topics-driven by emotional, contextual, or associative cues-a phenomenon known as spontaneous thought in neuroscience. In contrast, self-attention-based models rely on structured patterns over their inputs to predict each next token, lacking spontaneity. Motivated by this distinction, we characterize spontaneous topic changes in self-attention architectures and reveal divergences from spontaneous human thought. First, we establish theoretical results under a simplified, singlelayer self-attention model with suitable conditions by defining a topic as a set of Token Priority Graphs (TPGs). Specifically, we demonstrate that (1) the model maintains the priority order of tokens related to the input topic, (2) a spontaneous topic change can occur only if lower-priority tokens outnumber all higher-priority tokens of the input topic, and (3) unlike human cognition, the longer context length or the more ambiguous input topic does not increase the likelihood of spontaneous change. Second, we empirically validate that the effect of input length or topic ambiguity persists in modern, state-of-the-art LLMs, underscoring a fundamental disparity between human cognition and AI behavior in the context of spontaneous topic changes. To the best of our knowledge, no prior work has explored these questions with a focus so closely aligned to human thought.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1dc56cdd-a8a1-4eb4-bd6a-78996d1112b1Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
Related papers
- Mind the Gap: The Divergence Between Human and LLM-Generated TasksYi-Long Lu, Jiajun Song, Chunhui Zhang, Wei WangAAAI 2026
- Distributional Associations vs In-Context Reasoning: A Study of Feed-forward and Attention LayersLei Chen, Joan Bruna, Alberto BiettiICLR 2025
- Do LLMs dream of elephants (when told not to)? Latent concept association and associative memory in transformersYibo Jiang, Goutham Rajendran, Pradeep Ravikumar, Bryon AragamNeurIPS 2024 · 19 citations
- Prediction Hubs are Context-Informed Frequent Tokens in LLMsBeatrix Miranda Ginn Nielsen, Iuri Macocco, Marco BaroniACL 2025 · 2 citations
- From Noise to Narrative: Tracing the Origins of Hallucinations in TransformersPraneet Suresh, Jack Stanley, Sonia Joseph, Luca Scimeca et al.NeurIPS 2025 · 6 citations
