KBQA-R1: Reinforcing Large Language Models for Knowledge Base Question Answering
Xin Sun, Zhongqi Chen, Xing Zheng, Bowen Song, Qiang Liu, Shu Wu, Zilei Wang, Weiqiang Wang, Liang Wang
Abstract
Knowledge Base Question Answering (KBQA) challenges models to bridge the gap between natural language and strict knowledge graph schemas by generating executable logical forms. While Large Language Models (LLMs) have advanced this field, current approaches often struggle with a dichotomy of failure: they either generate hallucinated queries without verifying schema existence or exhibit rigid, template-based reasoning that mimics synthesized traces without true comprehension of the environment. To address these limitations, we present KBQA-R1, a framework that shifts the paradigm from text imitation to interaction optimization via Reinforcement Learning. Treating KBQA as a multi-turn decision process, our model learns to autonomously navigate the knowledge base using a structured action space, refining its reasoning strategies based on concrete execution feedback rather than static supervision. Furthermore, we introduce Referenced Rejection Sampling (RRS), a data synthesis method that resolves cold-start challenges by strictly aligning reasoning traces with ground-truth action sequences. Extensive experiments on WebQSP, GrailQA, and GraphQuestions demonstrate that KBQA-R1 achieves stateof-the-art performance. Code and project page are available at https://github. com/sunxin000/KBQA-R1 and https:// sunxin000.github.io/KBQA-R1/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 32e784e1-eb2d-4783-9bc8-534ed1fb1f0fBuilds on30
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
Related papers
- Generating then Refining for Reliable Knowledge Base Question AnsweringJianqi Gao, Hang Yu, Jian Cao, Ranran Bu et al.ACL 2026
- Temp-R1: A Unified Autonomous Agent for Complex Temporal KGQA via Reverse Curriculum Reinforcement LearningZhaoyan Gong, Zhiqiang Liu, Songze Li, Xiaoke Guo et al.ACL 2026 · 7 citations
- iQUEST: An Iterative Question-Guided Framework for Knowledge Base Question AnsweringShuai Wang, Yinan YuACL 2025 · 12 citations
- Plan Then Retrieve: Reinforcement Learning-Guided Complex Reasoning over Knowledge GraphsYanlin Song, Ben Liu, Víctor Gutiérrez-Basulto, Zhiwei Hu et al.WWW 2026
- RNG-KBQA: Generation Augmented Iterative Ranking for Knowledge Base Question AnsweringXi Ye, Semih Yavuz, Kazuma Hashimoto, Yingbo Zhou et al.ACL 2022 · 203 citations
