KBQA-R1: Reinforcing Large Language Models for Knowledge Base Question Answering
Xin Sun, Zhongqi Chen, Xing Zheng, Bowen Song, Qiang Liu, Shu Wu, Zilei Wang, Weiqiang Wang, Liang Wang
摘要
Knowledge Base Question Answering (KBQA) challenges models to bridge the gap between natural language and strict knowledge graph schemas by generating executable logical forms. While Large Language Models (LLMs) have advanced this field, current approaches often struggle with a dichotomy of failure: they either generate hallucinated queries without verifying schema existence or exhibit rigid, template-based reasoning that mimics synthesized traces without true comprehension of the environment. To address these limitations, we present KBQA-R1, a framework that shifts the paradigm from text imitation to interaction optimization via Reinforcement Learning. Treating KBQA as a multi-turn decision process, our model learns to autonomously navigate the knowledge base using a structured action space, refining its reasoning strategies based on concrete execution feedback rather than static supervision. Furthermore, we introduce Referenced Rejection Sampling (RRS), a data synthesis method that resolves cold-start challenges by strictly aligning reasoning traces with ground-truth action sequences. Extensive experiments on WebQSP, GrailQA, and GraphQuestions demonstrate that KBQA-R1 achieves stateof-the-art performance. Code and project page are available at https://github. com/sunxin000/KBQA-R1 and https:// sunxin000.github.io/KBQA-R1/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper30
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
相关 Paper
- Generating then Refining for Reliable Knowledge Base Question AnsweringJianqi Gao, Hang Yu, Jian Cao, Ranran Bu 等ACL 2026
- Temp-R1: A Unified Autonomous Agent for Complex Temporal KGQA via Reverse Curriculum Reinforcement LearningZhaoyan Gong, Zhiqiang Liu, Songze Li, Xiaoke Guo 等ACL 2026 · 被引用 7 次
- iQUEST: An Iterative Question-Guided Framework for Knowledge Base Question AnsweringShuai Wang, Yinan YuACL 2025 · 被引用 12 次
- Plan Then Retrieve: Reinforcement Learning-Guided Complex Reasoning over Knowledge GraphsYanlin Song, Ben Liu, Víctor Gutiérrez-Basulto, Zhiwei Hu 等WWW 2026
- RNG-KBQA: Generation Augmented Iterative Ranking for Knowledge Base Question AnsweringXi Ye, Semih Yavuz, Kazuma Hashimoto, Yingbo Zhou 等ACL 2022 · 被引用 203 次
