Step Back to Leap Forward: Self-Backtracking for Symbolic Reasoning and Planning in Language Models
Xiao-Wen Yang, Xuan-Yi Zhu, Dingchu Zhang, Wen-Da Wei, Jie-Jing Shao, Zhi Zhou, Lan-Zhe Guo, Yufeng Li
Abstract
Although autoregressive language models demonstrated remarkable performance across various tasks, their effectiveness in symbolic reasoning and decision-making scenarios remains constrained. Recent research indicates that training language models to emulate symbolic search algorithms (e.g. depth-first search or A* algorithm) can yield strong improvements in their symbolic reasoning and planning capabilities. However, existing methods only achieve superficial imitation of symbolic search trajectories, as their generation processes lack explicit backtracking mechanisms. This limitation prevents models from truly mastering symbolic search, often resulting in rigid and redundant outputs with poor solution quality. To address this issue, we propose a self-backtracking mechanism that enables LLMs to autonomously determine when to backtrack through specialized training, effectively utilizing this capability to scale during inference. By introducing a self-improvement strategy, the model can further refine its search process into optimal solution generation, improving problem-solving efficiency. Empirical evaluations demonstrate that our method boosts LLMs' reasoning on the Countdown task by 40% over optimal-path supervised fine-tuning (SFT) and improves both performance and efficiency on the Maze Navigation task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Non-Monotonic Autoregressive Sequence ModelTianyi MA, Yiyue Qian, Yiyang Li, Zehong Wang et al.ICML 2026
- Open-World LLM Logical ReasoningYe Mo, Chuan Zhou, Fengxiang Cheng, Jialin Yu et al.ICML 2026
Builds on15
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky et al.ICML 2024 · 973 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
Related papers
- Enhancing Logical Reasoning in Language Models via Symbolically-Guided Monte Carlo Process SupervisionXingwei Tan, Marco Valentino, Mahmud Elahi Akhter, Maria Liakata et al.EMNLP 2025 · 7 citations
- Learning to Better Search with Language Models via Guided Reinforced Self-TrainingSeungyong Moon, Bumsoo Park, Hyun Oh SongNeurIPS 2025 · 3 citations
- AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-PlayRan Xu, Yuchen Zhuang, Zihan Dong, Ruiyu Wang et al.NeurIPS 2025 · 10 citations
- Toward Self-Improvement of LLMs via Imagination, Searching, and CriticizingYe Tian, Baolin Peng, Linfeng Song, Lifeng Jin et al.NeurIPS 2024 · 162 citations
- Self-CriTeach: LLM Self-Teaching and Self-Critiquing for Improving Robotic Planning via Automated Domain GenerationJinbang Huang, Zhiyuan Li, Yuanzhao Hu, Zhanguang Zhang et al.ICML 2026 · 3 citations
