LEDOM: Reverse Language Model
Xunjian Yin, Sitao Cheng, Yuxi Xie, Xinyu Hu, Li Lin, Xinyi Wang, Liangming Pan, William Yang Wang, Xiaojun Wan
摘要
Autoregressive language models are trained exclusively left-to-right. We explore the complementary factorization, training right-to-left at scale, and ask what reasoning patterns emerge when a model conditions on future context to predict the past. We train LEDOM, an open-source purely reverse autoregressive language model (2B/7B parameters, 435B tokens), and find it develops capabilities distinct from forward models, including abductive inference, question synthesis, and natural resolution of the reversal curse. We then explore one application of the reverse model: combining forward likelihood with reverse posterior through noisy channel duality. We propose Reverse Reward, which reranks forward outputs using reverse posterior estimates, and prove that bidirectional scoring penalizes hallucinated reasoning chains whose backward reconstruction degrades. Reverse Reward yields gains of up to 6.6% on AIME 2024 and 15% on AMC 2023 across multiple strong baselines. We release all models, code, and data here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- Language Model InversionJohn X. Morris, Wenting Zhao, Justin T. Chiu, Vitaly Shmatikov 等ICLR 2024 · 被引用 6 次
- Time-Reversal Provides Unsupervised Feedback to LLMsYerram Varun, Rahul Madhavan, Sravanti Addepalli, Arun Suggala 等NeurIPS 2024 · 被引用 5 次
- OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction DataShubham Toshniwal, Wei Du, Ivan Moshkov, Branislav Kisacanin 等ICLR 2025
相关 Paper
- Towards a Theoretical Understanding of the 'Reversal Curse' via Training DynamicsHanlin Zhu, Baihe Huang, Shaolun Zhang, Michael I. Jordan 等NeurIPS 2024 · 被引用 37 次
- Breaking the Reversal Curse in Autoregressive Language Models via Identity BridgeXutao Ma, Yixiao Huang, Hanlin Zhu, Somayeh SojoudiICML 2026 · 被引用 2 次
- The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and MoreOuail Kitouni, Niklas Nolte, Adina Williams, Michael Rabbat 等NeurIPS 2024 · 被引用 29 次
- Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement LearningZhiheng Xi, Wenxiang Chen, Boyang Hong, Senjie Jin 等ICML 2024 · 被引用 70 次
- Library-Like Behavior In Language Models is Enhanced by Self-Referencing Causal CyclesMunachiso Nwadike, Zangir Iklassov, Toluwani Aremu, Tatsuya Hiraoka 等ACL 2025
