Delving into the Reversal Curse: How Far Can Large Language Models Generalize?
Zhengkai Lin, Zhihang Fu, Kai Liu, Liang Xie, Binbin Lin, Wenxiao Wang, Deng Cai, Yue Wu, Jieping Ye
Abstract
While large language models (LLMs) showcase unprecedented capabilities, they also exhibit certain inherent limitations when facing seemingly trivial tasks. A prime example is the recently debated"reversal curse", which surfaces when models, having been trained on the fact"A is B", struggle to generalize this knowledge to infer that"B is A". In this paper, we examine the manifestation of the reversal curse across various tasks and delve into both the generalization abilities and the problem-solving mechanisms of LLMs. This investigation leads to a series of significant insights: (1) LLMs are able to generalize to"B is A"when both A and B are presented in the context as in the case of a multiple-choice question. (2) This generalization ability is highly correlated to the structure of the fact"A is B"in the training documents. For example, this generalization only applies to biographies structured in"[Name] is [Description]"but not to"[Description] is [Name]". (3) We propose and verify the hypothesis that LLMs possess an inherent bias in fact recalling during knowledge application, which explains and underscores the importance of the document structure to successful learning. (4) The negative impact of this bias on the downstream performance of LLMs can hardly be mitigated through training alone. These findings offer a novel perspective on interpreting LLMs' generalization through their intrinsic mechanisms and provide insights for developing more effective learning methods. Our code and data are available at https://github.com/alibaba/thinking_bias.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Is the Reversal Curse a Binding Problem? Uncovering Limitations of Transformers from a Basic Generalization FailureBoshi Wang, Huan SunICLR 2026 · 16 citations
- Deep sequence models tend to memorize geometrically; it is unclear whyShahriar Noroozizadeh, Vaishnavh Nagarajan, Elan Rosenfeld, Sanjiv KumarICML 2026 · 11 citations
- Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric FactualityNitay Calderon, Eyal Ben-David, Zorik Gekhman, Eran Ofek et al.ICML 2026 · 8 citations
- Breaking the Reversal Curse in Autoregressive Language Models via Identity BridgeXutao Ma, Yixiao Huang, Hanlin Zhu, Somayeh SojoudiICML 2026 · 2 citations
- Memorization, Emergence, and Explaining Reversal Failures: A Controlled Study of Relational Semantics in LLMsYihua Zhu, Qianying Liu, Jiaxin Wang, Fei Cheng et al.ACL 2026 · 1 citation
Builds on22
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
Related papers
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"Lukas Berglund, Meg Tong, Maximilian Kaufmann, Mikita Balesni et al.ICLR 2024 · 462 citations
- An Analysis and Mitigation of the Reversal CurseAng Lv, Kaiyi Zhang, Shufang Xie, Quan Tu et al.EMNLP 2024
- Library-Like Behavior In Language Models is Enhanced by Self-Referencing Causal CyclesMunachiso Nwadike, Zangir Iklassov, Toluwani Aremu, Tatsuya Hiraoka et al.ACL 2025
- The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and MoreOuail Kitouni, Niklas Nolte, Adina Williams, Michael Rabbat et al.NeurIPS 2024 · 29 citations
- Bilinear representation mitigates reversal curse and enables consistent model editingDong-Kyum Kim, Minsung Kim, Jea Kwon, Nakyeong Yang et al.ICLR 2026 · 1 citation
