Flow of Reasoning: Training LLMs for Divergent Reasoning with Minimal Examples
Fangxu Yu, Lai Jiang, Haoqiang Kang, Shibo Hao, Lianhui Qin
Abstract
The ability to generate diverse solutions to a given problem is a hallmark of human creativity. This divergent reasoning is also crucial for machines, enhancing their robustness and enabling them to assist humans in many applications such as scientific discovery. However, existing approaches to multi-step reasoning with large language models (LLMs) have mostly focused only on reasoning accuracy, without further discovering more diverse valid solutions. For example, supervised fine-tuning improves reasoning quality but requires vast labeled data, while reward-maximizing reinforcement learning finds top-reward solutions while neglecting the solution diversity. To fill this gap, we propose Flow of Reasoning (FOR), an efficient diversity-seeking LLM finetuning method aimed at improving reasoning quality and diversity with minimal data. FOR formulates multi-step LLM reasoning as a Markovian flow on a DAG-structured reasoning graph. This formulation allows us to incorporate and adapt principled GFlowNet approaches, for finetuning LLMs to sample divergent paths with probabilities proportional to the (unnormalized) reward of target problems. Extensive experiments show that, with limited training examples (e.g., 15 examples), FOR enables the discovery of diverse, creative, high-quality solutions, greatly outperforming a wide range of existing inference and training methods across six challenging reasoning tasks, including BlocksWorld (embodied reasoning), Game24 (math puzzle solving), Rubik's Cube (spatial reasoning), 1D-ARC (abstraction reasoning), GSM8k (math reasoning), and Pron-toQA (logical reasoning). Code is available at https://github.com/Yu-Fangxu/FoR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 06b896a6-3ced-44fd-8588-9bb2b610f558Cited by top-tier papers6
- FlowRL: Matching Reward Distributions for LLM ReasoningXuekai Zhu, Daixuan Cheng, Dinghuai Zhang, Hengli Li et al.ICLR 2026 · 41 citations
- LaDiR: Latent Diffusion Enhances LLMs for Text ReasoningHaoqiang Kang, Yizhe Zhang, Nikki Lijing Kuang, Nicklas Majamaki et al.ICLR 2026 · 25 citations
- SeLaR: Selective Latent Reasoning in Large Language ModelsRenyu Fu, Guibo LuoACL 2026 · 2 citations
- Flow of Spans: Generalizing Language Models to Dynamic Span-Vocabulary via GFlowNetsBo Xue, Yunchong Song, Fanghao Shao, Xuekai Zhu et al.ICLR 2026
- Beyond Normalization: Rethinking the Partition Function as a Difficulty Scheduler for RLVRDohyung Kim, Minbeom Kim, Jeonghye Kim, Lee Sangmook et al.ICML 2026
Builds on58
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
Related papers
- Amortizing intractable inference in large language modelsEdward J. Hu, Moksh Jain, Eric Elmoznino, Younesse Kaddar et al.ICLR 2024 · 91 citations
- GFlowVLM: Enhancing Multi-step Reasoning in Vision-Language Models with Generative Flow NetworksHaoqiang Kang, Enna Sachdeva, Piyush Gupta, Sangjae Bae et al.CVPR 2025
- GPO: Learning from Critical Steps to Improve LLM ReasoningJiahao Yu, Zelei Cheng, Xian Wu, Xinyu XingNeurIPS 2025 · 10 citations
- Training Large Language Models To Reason In Parallel With Global Forking TokensSheng Jia, Xiao Wang, Shiva Prasad KasiviswanathanICLR 2026 · 7 citations
- Reasoning Quality Emerges Early: Data Curation for Reasoning ModelsHongyi Jin, Wenhan Yang, Meysam Ghaffari, Carlos Morato et al.ICML 2026
