Step Back to Leap Forward: Self-Backtracking for Symbolic Reasoning and Planning in Language Models
Xiao-Wen Yang, Xuan-Yi Zhu, Dingchu Zhang, Wen-Da Wei, Jie-Jing Shao, Zhi Zhou, Lan-Zhe Guo, Yufeng Li
摘要
Although autoregressive language models demonstrated remarkable performance across various tasks, their effectiveness in symbolic reasoning and decision-making scenarios remains constrained. Recent research indicates that training language models to emulate symbolic search algorithms (e.g. depth-first search or A* algorithm) can yield strong improvements in their symbolic reasoning and planning capabilities. However, existing methods only achieve superficial imitation of symbolic search trajectories, as their generation processes lack explicit backtracking mechanisms. This limitation prevents models from truly mastering symbolic search, often resulting in rigid and redundant outputs with poor solution quality. To address this issue, we propose a self-backtracking mechanism that enables LLMs to autonomously determine when to backtrack through specialized training, effectively utilizing this capability to scale during inference. By introducing a self-improvement strategy, the model can further refine its search process into optimal solution generation, improving problem-solving efficiency. Empirical evaluations demonstrate that our method boosts LLMs' reasoning on the Countdown task by 40% over optimal-path supervised fine-tuning (SFT) and improves both performance and efficiency on the Maze Navigation task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Non-Monotonic Autoregressive Sequence ModelTianyi MA, Yiyue Qian, Yiyang Li, Zehong Wang 等ICML 2026
- Open-World LLM Logical ReasoningYe Mo, Chuan Zhou, Fengxiang Cheng, Jialin Yu 等ICML 2026
它引用的顶会 Paper15
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky 等ICML 2024 · 被引用 973 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
相关 Paper
- Enhancing Logical Reasoning in Language Models via Symbolically-Guided Monte Carlo Process SupervisionXingwei Tan, Marco Valentino, Mahmud Elahi Akhter, Maria Liakata 等EMNLP 2025 · 被引用 7 次
- Learning to Better Search with Language Models via Guided Reinforced Self-TrainingSeungyong Moon, Bumsoo Park, Hyun Oh SongNeurIPS 2025 · 被引用 3 次
- AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-PlayRan Xu, Yuchen Zhuang, Zihan Dong, Ruiyu Wang 等NeurIPS 2025 · 被引用 10 次
- Toward Self-Improvement of LLMs via Imagination, Searching, and CriticizingYe Tian, Baolin Peng, Linfeng Song, Lifeng Jin 等NeurIPS 2024 · 被引用 162 次
- Self-CriTeach: LLM Self-Teaching and Self-Critiquing for Improving Robotic Planning via Automated Domain GenerationJinbang Huang, Zhiyuan Li, Yuanzhao Hu, Zhanguang Zhang 等ICML 2026 · 被引用 3 次
