RAG-on-a-Diet: A Reinforcement Learning-Based Dynamic Resource Optimization Framework for RAG
Hongwen Ding, Yizheng Zhao
摘要
Retrieval-Augmented Generation (RAG) has become the backbone of knowledge-intensive multi-hop question answering, yet routing every sub-query through a frontier model turns every hop into a cost multiplier and makes realworld deployment prohibitively expensive. Existing remedies either fix the retrieval schedule, route once at the query level, or lack a principled stopping rule, leaving a critical gap: no framework adapts, hop by hop, to how a trajectory actually unfolds. We introduce RAG-on-a-Diet, a lightweight reinforcementlearning agent that treats each reasoning hop as an independent decision and selects the smallest model (Qwen3-4B, Qwen3-30B, or DS-R1-671B) sufficient for it, guided by entityand confidence-aware features. Trained via behavior cloning followed by PPO under a fivecomponent cost-aware reward (final, cumulative, step-wise, cost, balance) and coupled with an explicit two-tier termination policy (5-hop cap plus a τ = 0.3 confidence gate), the agent carves a Pareto-optimal efficiency frontier. On HotpotQA it cuts Monetary Inference Cost by 60.07% against IRCoT with only a 3.7% F1 drop; it matches Adaptive-RAG's F1 at 37.30% lower cost; and it attains up to 2.33× higher Quality-per-Monetary-Cost. Consistent gains on MuSiQue, 2WikiMultiHopQA, CRAG, and Bamboogle confirm strong out-of-distribution robustness, setting a new paradigm for finegrained resource control in multi-hop RAG.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat 等ICML 2020 · 被引用 2,937 次
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil 等ICLR 2024 · 被引用 1,798 次
相关 Paper
- CompactRAG: Reducing LLM Calls and Token Overhead in Multi-Hop Question AnsweringHao Yang, Zhiyu Yang, Xupeng Zhang, Wei Wei 等WWW 2026
- FrugalRAG: Less is More in RL Finetuning for Multi-hop Question AnsweringAbhinav Java, Srivathsan Koundinyan, Nagarajan Natarajan, Amit SharmaICLR 2026 · 被引用 2 次
- S2G-RAG: Structured Sufficiency and Gap Judging for Iterative Retrieval-Augmented QAMinghan Li, Junjie Zou, Xinxuan Lv, Chao Zhang 等ACL 2026 · 被引用 1 次
- HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented GenerationPeilin Wu, Mian Zhang, Kun Wan, Wentian Zhao 等ICLR 2026 · 被引用 13 次
- LiR3AG: A Lightweight Rerank Reasoning Strategy Framework for Retrieval-Augmented GenerationGuo Chen, Junjie Huang, Huaijin Xie, Fei Sun 等AAAI 2026
