Breaking the Reversal Curse in Autoregressive Language Models via Identity Bridge
Xutao Ma, Yixiao Huang, Hanlin Zhu, Somayeh Sojoudi
Abstract
Autoregressive large language models (LLMs) have achieved remarkable success in many complex tasks, yet they can still fail in very simple logical reasoning such as the "reversal curse" --- when trained on forward knowledge data of the form "" (e.g., Alice's husband is Bob), the model is unable to deduce the reversal knowledge "" (e.g., Bob's wife is Alice) during test. Extensive prior research suggests that this failure is an inherent, fundamental limit of autoregressive causal LLMs, indicating that these models tend to memorize factual-level knowledge rather than capture higher-level rules. In this paper, we challenge this view by showing that this seemingly fundamental limit can be mitigated by slightly tweaking the training data with a simple regularization data recipe called the Identity Bridge of the form "" (e.g., The name of Alice is Alice). Theoretically, we prove that under this recipe, even a one-layer transformer can break the reversal curse by analyzing the implicit bias of gradient descent. Empirically, we show that a 1B pretrained language model finetuned with the proposed data recipe achieves a 40% success rate on reversal tasks, in stark contrast to a near-zero success rate when trained solely on forward-knowledge data. Our work provides a novel theoretical foundation for the reversal curse and offers a principled, low-cost path to encouraging LLMs to learn higher-level rules from data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c50d7e2a-53c6-4748-8292-98ff9e6bd40cCited by top-tier papers1
Ask how each one uses itBuilds on24
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang et al.NeurIPS 2025 · 949 citations
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"Lukas Berglund, Meg Tong, Maximilian Kaufmann, Mikita Balesni et al.ICLR 2024 · 462 citations
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-AttentionArvind V. Mahankali, Tatsunori Hashimoto, Tengyu MaICLR 2024 · 160 citations
- Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer TransformerYuandong Tian, Yiping Wang, Beidi Chen, Simon S. DuNeurIPS 2023 · 125 citations
Related papers
- Towards a Theoretical Understanding of the 'Reversal Curse' via Training DynamicsHanlin Zhu, Baihe Huang, Shaolun Zhang, Michael I. Jordan et al.NeurIPS 2024 · 37 citations
- An Analysis and Mitigation of the Reversal CurseAng Lv, Kaiyi Zhang, Shufang Xie, Quan Tu et al.EMNLP 2024
- Delving into the Reversal Curse: How Far Can Large Language Models Generalize?Zhengkai Lin, Zhihang Fu, Kai Liu, Liang Xie et al.NeurIPS 2024 · 12 citations
- Bilinear representation mitigates reversal curse and enables consistent model editingDong-Kyum Kim, Minsung Kim, Jea Kwon, Nakyeong Yang et al.ICLR 2026 · 1 citation
- Library-Like Behavior In Language Models is Enhanced by Self-Referencing Causal CyclesMunachiso Nwadike, Zangir Iklassov, Toluwani Aremu, Tatsuya Hiraoka et al.ACL 2025
