Interact-RAG: Reason and Interact with the Corpus, Beyond Black-Box Retrieval
Yulong Hui, Chao Chen, Zhihang Fu, Yihao Liu, Jieping Ye, Huanchen Zhang
Abstract
Retrieval-Augmented Generation (RAG) has significantly enhanced LLMs by incorporating external information. However, prevailing agentic RAG approaches are constrained by a critical limitation: they treat the retrieval process as a black-box querying operation. This confines agents' actions to query issuing, hindering its ability to tackle complex information-seeking tasks. To address this, we introduce Interact-RAG, a new paradigm that elevates the LLM agent from a passive query issuer into an active manipulator of the retrieval process. We dismantle the black-box with a Corpus Interaction Engine, equipping the agent with a set of action primitives for fine-grained control over information retrieval. To further empower the agent on the entire RAG pipeline, we first develop a reasoning-enhanced workflow, which enables both zero-shot execution and the synthesis of interaction trajectories. We then leverage this synthetic data to train a fully autonomous end-to-end agent via Supervised Fine-Tuning (SFT), followed by refinement with Reinforcement Learning (RL). Extensive experiments across six benchmarks demonstrate that Interact-RAG significantly outperforms other advanced methods, validating the efficacy of our reasoning-interaction strategy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- ScaleDoc: Scaling LLM-based Predicates over Large Document CollectionsHengrui Zhang, Yulong Hui, Yihao Liu, Huanchen ZhangSIGMOD 2026 · 2 citations
- -Reader: Dual Evolving Graphs for Multimodal Document QAYaxin Du, Junru Song, Yifan Zhou, Cheng Wang et al.ICML 2026 · 1 citation
- Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development ResearchNimisha Karnatak, Mohamad Chatila, Daniel Alejandro Pinzón Hernández, Reza Yazdanfar et al.CHI 2026 · 1 citation
Builds on13
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Active Retrieval Augmented GenerationZhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun et al.EMNLP 2023 · 315 citations
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das et al.ACL 2023 · 233 citations
- Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step QuestionsHarsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish SabharwalACL 2023 · 187 citations
Related papers
- Optimizing Retrieval for RAG via Reinforcement LearningJiawei Zhou, Lei ChenNeurIPS 2025 · 1 citation
- Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement LearningWenlin Zhang, Xiangyang Li, Kuicai Dong, Yichao Wang et al.NeurIPS 2025 · 85 citations
- InstructRAG: Instructing Retrieval-Augmented Generation via Self-Synthesized RationalesZhepei Wei, Wei-Lin Chen, Yu MengICLR 2025
- s3: You Don't Need That Much Data to Train a Search Agent via RLPengcheng Jiang, Xueqiang Xu, Jiacheng Lin, Jinfeng Xiao et al.EMNLP 2025
- Graph-R1: Towards Agentic GraphRAG Framework via End-to-end Reinforcement LearningHaoran Luo, Haihong E, Guanting Chen, Qika Lin et al.ICML 2026 · 50 citations
