From Complex to Atomic: Enhancing Augmented Generation via Knowledge-Aware Dual Rewriting and Reasoning
Jinyu Wang, Jingjing Fu, Rui Wang, Lei Song, Jiang Bian
Abstract
Recent advancements in Retrieval-Augmented Generation (RAG) systems have significantly enhanced the capabilities of large language models (LLMs) by incorporating external knowledge retrieval. However, the sole reliance on retrieval is often inadequate for mining deep, specialized knowledge and performing the logical reasoning necessary to tackle domain-specific complex questions. To address these challenges, we present an approach, which is designed to extract, comprehend, and utilize specialized knowledge in an atomic manner while simultaneously constructing a coherent rationale. At the heart of our approach lie four pivotal components: a knowledge atomizer that extracts atomic tags from raw data, a query proposer that generates subsequent questions to facilitate the original inquiry, an atomic retriever that locates knowledge based on atomic knowledge alignments, and an atomic selector that determines which atomic tag and chunk pair to query, guided by the retrieved information. Through this approach, we implement a knowledge-aware task decomposition strategy that iteratively builds the rationale in alignment with the initial question and the acquired knowledge. We conduct comprehensive experiments to demonstrate the efficacy of our approach across various benchmarks, particularly those requiring multihop reasoning steps. A substantial performance improvement of up to +10.1 (20.4%) over the second-best method underscores the potential of the approach in complex, knowledgeintensive applications. The code is publicly available at https://github.com/microsoft/PIKE-RAG .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
- Active Retrieval Augmented GenerationZhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun et al.EMNLP 2023 · 315 citations
- Knowledge Graph Prompting for Multi-Document Question AnsweringYu Wang, Nedim Lipka, Ryan A. Rossi, Alexa F. Siu et al.AAAI 2024 · 290 citations
Related papers
- AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge ReasoningAmy Xin, Jinxin Liu, Zijun Yao, Zhicheng Lee et al.KDD 2025 · 1 citation
- DeepRAG: Thinking to Retrieve Step by Step for Large Language ModelsXinyan Guan, Jiali Zeng, Fandong Meng, Chunlei Xin et al.ICLR 2026 · 30 citations
- RAG+: Enhancing Retrieval-Augmented Generation with Application-Aware ReasoningYu Wang, Shiwan Zhao, Zhihu Wang, Ming Fan et al.EMNLP 2025 · 3 citations
- Clue-RAG: Towards Accurate and Cost-Efficient Graph-Based RAG Via Multi-Partite Graph-Based IndexYaodong Su, Yixiang Fang, Yingli Zhou, Chuanhui YangICDE 2026
- T-GRAG: A Dynamic GraphRAG Framework for Resolving Temporal Conflicts and Redundancy in Knowledge RetrievalDong Li, Yichen Niu, Ying Ai, Xiang Zou et al.ACM MM 2025 · 11 citations
