Explore-on-Graph: Incentivizing Autonomous Exploration of Large Language Models on Knowledge Graphs with Path-refined Reward Modeling
Shiqi Yan, Yubo Chen, Ruiqi Zhou, Zhengxi Yao, Shuai Chen, Tianyi Zhang, Shijie Zhang, Wei Qiang Zhang, Yongfeng Huang, Haixin Duan, Yunqi Zhang
Abstract
The reasoning process of Large Language Models (LLMs) is often plagued by hallucinations and missing facts in question-answering tasks. A promising solution is to ground LLMs' answers in verifiable knowledge sources, such as Knowledge Graphs (KGs). Prevailing KG-enhanced methods typically constrained LLM reasoning either by enforcing rules during generation or by imitating paths from a fixed set of demonstrations. However, they naturally confined the reasoning patterns of LLMs within the scope of prior experience or fine-tuning data, limiting their generalizability to out-of-distribution graph reasoning problems. To tackle this problem, in this paper, we propose Explore-on-Graph (EoG), a novel framework that encourages LLMs to autonomously explore a more diverse reasoning space on KGs. To incentivize exploration and discovery of novel reasoning paths, we propose to introduce reinforcement learning during training, whose reward is the correctness of the reasoning paths' final answers. To enhance the efficiency and meaningfulness of the exploration, we propose to incorporate path information as additional reward signals to refine the exploration process and reduce futile efforts. Extensive experiments on five KGQA benchmark datasets demonstrate that, to the best of our knowledge, our method achieves state-of-the-art performance, outperforming not only open-source but also even closed-source LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d3d8041-8e74-482c-8f66-aa1d97d5695eCited by top-tier papers2
- Autonomous Knowledge Graph Exploration with Adaptive Breadth-Depth RetrievalJoaquín Polonuer, Lucas Vittor, Iñaki Arango, Ayush Noori et al.ACL 2026 · 1 citation
- SHAPE: Unifying Safety, Helpfulness and Pedagogy for Educational LLMsSihang Zhao, Kangrui Yu, Youliang Yuan, Pinjia He et al.ACL 2026
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Reasoning on Graphs: Faithful and Interpretable Large Language Model ReasoningLinhao Luo, Yuan-Fang Li, Gholamreza Haffari, Shirui PanICLR 2024 · 499 citations
- Beyond I.I.D.: Three Levels of Generalization for Question Answering on Knowledge BasesYu Gu, Sue Kase, Michelle Vanni, Brian M. Sadler et al.WWW 2021 · 304 citations
- Subgraph Retrieval Enhanced Model for Multi-hop Knowledge Base Question AnsweringJing Zhang, Xiaokang Zhang, Jifan Yu, Jian Tang et al.ACL 2022 · 221 citations
Related papers
- Enrich-on-Graph: Query-Graph Alignment for Complex Reasoning with LLM EnrichingSongze Li, Zhiqiang Liu, Zhengke Gui, Huajun Chen et al.EMNLP 2025 · 6 citations
- Plan-Answer-Refine-on-Graph: Structured Planning and Self-Refinement for Large Language Model Reasoning on Knowledge GraphsYuxin Shi, Han Fu, Zhuo Li, Chenghao Liu et al.ICLR 2026
- Search-on-Graph: Iterative Informed Navigation for Large Language Model Reasoning on Knowledge GraphsJia Ao Sun, Hao Yu, Fabrizio Gotti, Fengran Mo et al.KDD 2026 · 8 citations
- Backjump-on-Graph: Empowering Large Language Models with Reinforced Retrospective Exploration for Agentic Knowledge Graph ReasoningYunqi Zhang, Shiqi Yan, Zhenzhao Yuan, Wenrui Liang et al.ICML 2026
- Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge GraphsLiyi Chen, Panrong Tong, Zhongming Jin, Ying Sun et al.NeurIPS 2024 · 160 citations
