Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers
Yixiao Huang, Hanlin Zhu, Tianyu Guo, Jiantao Jiao, Somayeh Sojoudi, Michael I. Jordan, Stuart J. Russell, Song Mei
Abstract
Large language models (LLMs) can acquire new knowledge through fine-tuning, but this process exhibits a puzzling duality: models can generalize remarkably from new facts, yet are also prone to hallucinating incorrect information. However, the reasons for this phenomenon remain poorly understood. In this work, we argue that both behaviors stem from a single mechanism known as out-of-context reasoning (OCR): the ability to deduce implications by associating concepts, even those without a causal link. Our experiments across five prominent LLMs confirm that OCR indeed drives both generalization and hallucination, depending on whether the associated concepts are causally related. To build a rigorous theoretical understanding of this phenomenon, we then formalize OCR as a synthetic factual recall task. We empirically show that a one-layer single-head attention-only transformer with factorized output and value matrices can learn to solve this task, while a model with combined weights cannot, highlighting the crucial role of matrix factorization. Our theoretical analysis shows that the OCR capability can be attributed to the implicit bias of gradient descent, which favors solutions that minimize the nuclear norm of the combined output-value matrix. This mathematical structure explains why the model learns to associate facts and implications with high sample efficiency, regardless of whether the correlation is causal or merely spurious. Ultimately, our work provides a theoretical foundation for understanding the OCR phenomenon, offering a new lens for analyzing and mitigating undesirable behaviors from knowledge injection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Emergence of Superposition: Unveiling the Training Dynamics of Chain of Continuous ThoughtHanlin Zhu, Shibo Hao, Zhiting Hu, Jiantao Jiao et al.ICLR 2026 · 21 citations
- Deep sequence models tend to memorize geometrically; it is unclear whyShahriar Noroozizadeh, Vaishnavh Nagarajan, Elan Rosenfeld, Sanjiv KumarICML 2026 · 11 citations
- Breaking the Reversal Curse in Autoregressive Language Models via Identity BridgeXutao Ma, Yixiao Huang, Hanlin Zhu, Somayeh SojoudiICML 2026 · 2 citations
- Why Deep Jacobian Spectra Separate: Depth-Induced Scaling and Singular-Vector AlignmentNathanaël Haas, François Gatine, Augustin Cosse, Zied BouraouiICML 2026 · 1 citation
- Learning to Recall with Transformers Beyond Orthogonal EmbeddingsNuri Mert Vural, Alberto Bietti, Mahdi Soltanolkotabi, Denny WuICLR 2026 · 1 citation
Builds on29
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"Lukas Berglund, Meg Tong, Maximilian Kaufmann, Mikita Balesni et al.ICLR 2024 · 462 citations
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das et al.ACL 2023 · 233 citations
- Birth of a Transformer: A Memory ViewpointAlberto Bietti, Vivien Cabannes, Diane Bouchacourt, Hervé Jégou et al.NeurIPS 2023 · 182 citations
- Implicit Regularization in Deep Learning May Not Be Explainable by NormsNoam Razin, Nadav CohenNeurIPS 2020 · 178 citations
Related papers
- Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training DataJohannes Treutlein, Dami Choi, Jan Betley, Samuel Marks et al.NeurIPS 2024
- When Do Hallucinations Arise? A Graph Perspective on the Evolution of Path Reuse and Path CompressionXinnan Dai, Kai Yang, cheng Luo, Shenglai Zeng et al.ICML 2026
- Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal et al.EMNLP 2024 · 53 citations
- Why LLMs Hallucinate on Structured Knowledge: A Mechanistic Analysis of Reasoning over Linearized RepresentationsShanghao Li, Jinda Han, Yibo Wang, Yuanjie Zhu et al.ACL 2026
- Enough Coin Flips Can Make LLMs Act BayesianRitwik Gupta, Rodolfo Corona, Jiaxin Ge, Eric Wang et al.ACL 2025
