Reversal of Thought: Enhancing Large Language Models with Preference-Guided Reverse Reasoning Warm-up
Jiahao Yuan, Dehui Du, Hao Zhang, Zixiang Di, Usman Naseem
摘要
Large language models (LLMs) have shown remarkable performance in reasoning tasks but face limitations in mathematical and complex logical reasoning. Existing methods to improve LLMs' logical capabilities either involve traceable or verifiable logical sequences that generate more reliable responses by constructing logical structures yet increase computational costs, or introduces rigid logic template rules, reducing flexibility. In this paper, we propose Reversal of Thought (RoT), a plug-and-play and cost-effective reasoning framework designed to enhance the logical reasoning abilities of LLMs during the warm-up phase prior to batch inference. RoT utilizes a Preference-Guided Reverse Reasoning warm-up strategy, which integrates logical symbols for pseudocode planning through meta-cognitive mechanisms and pairwise preference self-evaluation to generate task-specific prompts solely through demonstrations, aligning with LLMs' cognitive preferences shaped by RLHF. Through reverse reasoning, we utilize a Cognitive Preference Manager to assess knowledge boundaries and further expand LLMs' reasoning capabilities by aggregating solution logic for known tasks and stylistic templates for unknown tasks. Experiments across various tasks demonstrate that RoT surpasses existing baselines in both reasoning accuracy and efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-SteeringJinhe Bi, Yujun Wang, Haokun Chen, Xun Xiao 等ACL 2025
- The Geometry of Reasoning: Self-Evaluation via Layerwise Trajectory EvolutionJinhe Bi, Danqi Yan, Yifan Wang, Wenke Huang 等ICML 2026
它引用的顶会 Paper12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Graph of Thoughts: Solving Elaborate Problems with Large Language ModelsMaciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger 等AAAI 2024 · 被引用 1,292 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
相关 Paper
- RL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement LearningQianyue Hao, Sibo Li, Jian Yuan, Yong LiICLR 2026 · 被引用 20 次
- Retrieval-of-Thought: Efficient Reasoning via Reusing ThoughtsAmmar Ahmed, Azal Ahmad Khan, Ayaan Ahmad, Sheng Di 等ICLR 2026 · 被引用 13 次
- Adaption-of-Thought: Learning Question Difficulty Improves Large Language Models for ReasoningMayi Xu, Yongqi Li, Ke Sun, Tieyun QianEMNLP 2024 · 被引用 1 次
- Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent ReasoningYifan Wang, Shiyu Li, Peiming Li, Xiaochen Yang 等ACL 2026 · 被引用 14 次
- Faithful Logical Reasoning via Symbolic Chain-of-ThoughtJundong Xu, Hao Fei, Liangming Pan, Qian Liu 等ACL 2024
