Structure Guided Retrieval-Augmented Generation for Factual Queries
Miao Xie, Xiao Zhang, Yi Li, Chunli Lv
Abstract
Retrieval-Augmented Generation (RAG) has been proposed to mitigate hallucinations in large language models (LLMs), where generated outputs may be factually incorrect. However, existing RAG approaches predominantly rely on vector similarity for retrieval, which is prone to semantic noise and fails to ensure that generated responses fully satisfy the complex conditions specified by factual queries. To address this challenge, we introduce a novel research problem, named EXACT RETRIEVAL PROBLEM (ERP). To the best of our knowledge, this is the first problem formulation that explicitly incorporates structural information into RAG for factual questions to satisfy all query conditions. For this novel problem, we propose STRUCTURE GUIDED RETRIEVAL-AUGMENTED GENERATION (SG-RAG) 1 , which models the retrieval process as an embedding-based subgraph matching task, and uses the retrieved topological structures to guide the LLM to generate answers that meet all specified query conditions. To facilitate evaluation of ERP, we construct and publicly release EXACT RETRIEVAL QUESTION ANSWERING (ERQA) 2 , a large-scale dataset comprising 120,000 fact-oriented QA pairs, each involving complex conditions, spanning 20 diverse domains. The experimental results demonstrate that SG-RAG significantly outperforms strong baselines on ERQA, delivering absolute gains of 20.68-50.88 percentage points, corresponding to 34%-450% relative improvements across metrics, while maintaining reasonable computational overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b3257774-b370-4abb-8fb3-867d95b19d5fBuilds on10
- RAPTOR: Recursive Abstractive Processing for Tree-Organized RetrievalParth Sarthi, Salman Abdullah, Aditi Tuli, Shubh Khanna et al.ICLR 2024 · 460 citations
- HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language ModelsBernal Jimenez Gutierrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga et al.NeurIPS 2024 · 395 citations
- Precise Zero-Shot Dense Retrieval without Relevance LabelsLuyu Gao, Xueguang Ma, Jimmy Lin, Jamie CallanACL 2023 · 211 citations
- MedRAG: Enhancing Retrieval-augmented Generation with Knowledge Graph-Elicited Reasoning for Healthcare CopilotXuejiao Zhao, Siyan Liu, Su-Yin Yang, Chunyan MiaoWWW 2025 · 134 citations
- RapidMatch: A Holistic Approach to Subgraph Query ProcessingShixuan Sun, Xibo Sun, Yulin Che, Qiong Luo et al.VLDB 2021 · 105 citations
Related papers
- In-depth Analysis of Graph-based RAG in a Unified FrameworkYingli Zhou, Yaodong Su, Youran Sun, Shu Wang et al.VLDB 2025 · 48 citations
- RAGEval: Scenario Specific RAG Evaluation Dataset Generation FrameworkKunlun Zhu, Yifan Luo, Dingling Xu, Yukun Yan et al.ACL 2025 · 53 citations
- REAP: Enhancing RAG with Recursive Evaluation and Adaptive Planning for Multi-Hop Question AnsweringYijie Zhu, Haojie Zhou, Wanting Hong, Tailin Liu et al.AAAI 2026
- ArchRAG: Attributed Community-based Hierarchical Retrieval-Augmented GenerationShu Wang, Yixiang Fang, Yingli Zhou, Xilin Liu et al.AAAI 2026 · 23 citations
- You Don't Need Pre-Built Graphs for RAG: Retrieval Augmented Generation with Adaptive Reasoning StructuresShengyuan Chen, Chuang Zhou, Zheng Yuan, Qinggang Zhang et al.AAAI 2026 · 14 citations
