Flow-guided Direct Preference Optimization for Knowledge Graph Reasoning with Trees
Tiesunlong Shen, Rui Mao, Jin Wang, Xuejie Zhang, Erik Cambria
Abstract
Recent advancements in knowledge graph question answering (KGQA) have shown promise, yet existing methods often fail to align with human reasoning patterns that involve continuous reflection and refinement. This paper proposes FD-PORT (flow-guided direct preference optimization for knowledge graph reasoning with trees), a novel approach that combines Monte Carlo Tree Search (MCTS) with flow-guided direct preference optimization (FDPO) for KGQA tasks. MCTS simulates human-like reasoning by systematically exploring multiple inference paths in knowledge graphs, while FDPO transforms the search feedback into fine-grained training signals through flow balance conditions. Unlike traditional methods focusing on end-to-end training or sequence-level preferences, FD-PORT establishes flow consistency between any states along the reasoning chain, enabling robust multi-hop reasoning that adapts to local decisions and long-range dependencies. Experimental results on three benchmark datasets demonstrate that FD-PORT significantly outperforms state-of-the-art methods, achieving up to 50.6% improvements over GPT-4 on complex multi-hop reasoning tasks with a smaller open-source language model. The framework is advanced in maintaining diverse reasoning paths while ensuring answer quality, closely mirroring human problem-solving strategies.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers8
- MARS: Multi-Agent Adaptive Reasoning with Socratic Guidance for Automated Prompt OptimizationJian Zhang, Zhangqi Wang, Haiping Zhu, Kangda Cheng et al.AAAI 2026 · 9 citations
- From Stimuli to Minds: Enhancing Psychological Reasoning in LLMs via Bilateral Reinforcement LearningYichao Feng, Haoran Luo, Lang Feng, Shuai Zhao et al.AAAI 2026 · 4 citations
- LLMdoctor: Token-Level Flow-Guided Preference Optimization for Efficient Test-Time Alignment of Large Language ModelsTiesunlong Shen, Rui Mao, Jin Wang, Heming Sun et al.AAAI 2026 · 2 citations
- Step-GRPO: Enhancing Reasoning Quality and Efficiency via Structured PRM-Based Reinforcement LearningWeijie Li, Jin Wang, Liang-Chih Yu, Xuejie ZhangAAAI 2026 · 1 citation
- SAPO: Self-Adaptive Process Optimization Makes Small Reasoners StrongerKaiyuan Chen, Guangmin Zheng, Jin Wang, Xiaobing Zhou et al.AAAI 2026 · 1 citation
Related papers
- DAMR: Efficient and Adaptive Context-Aware Knowledge Graph Question Answering with LLM-Guided MCTSYingxu Wang, Shiqi Fan, Mengzhu Wang, Siyang Gao et al.ICLR 2026 · 7 citations
- Reinforcement Learning Enhanced Muti-hop Reasoning for Temporal Knowledge Question AnsweringWuzhenghong Wen, Chao Xue, Su Pan, Yuwei Sun et al.AAAI 2026
- Paths-over-Graph: Knowledge Graph Empowered Large Language Model ReasoningXingyu Tan, Xiaoyang Wang, Qing Liu, Xiwei Xu et al.WWW 2025 · 86 citations
- Ontology-Guided Reverse Thinking Makes Large Language Models Stronger on Knowledge Graph Question AnsweringRunxuan Liu, Bei Luo, Jiaqi Li, Baoxin Wang et al.ACL 2025 · 21 citations
- UniKGQA: Unified Retrieval and Reasoning for Solving Multi-hop Question Answering Over Knowledge GraphJinhao Jiang, Kun Zhou, Xin Zhao, Ji-Rong WenICLR 2023 · 28 citations
