RAG-on-a-Diet: A Reinforcement Learning-Based Dynamic Resource Optimization Framework for RAG
Hongwen Ding, Yizheng Zhao
Abstract
Retrieval-Augmented Generation (RAG) has become the backbone of knowledge-intensive multi-hop question answering, yet routing every sub-query through a frontier model turns every hop into a cost multiplier and makes realworld deployment prohibitively expensive. Existing remedies either fix the retrieval schedule, route once at the query level, or lack a principled stopping rule, leaving a critical gap: no framework adapts, hop by hop, to how a trajectory actually unfolds. We introduce RAG-on-a-Diet, a lightweight reinforcementlearning agent that treats each reasoning hop as an independent decision and selects the smallest model (Qwen3-4B, Qwen3-30B, or DS-R1-671B) sufficient for it, guided by entityand confidence-aware features. Trained via behavior cloning followed by PPO under a fivecomponent cost-aware reward (final, cumulative, step-wise, cost, balance) and coupled with an explicit two-tier termination policy (5-hop cap plus a τ = 0.3 confidence gate), the agent carves a Pareto-optimal efficiency frontier. On HotpotQA it cuts Monetary Inference Cost by 60.07% against IRCoT with only a 3.7% F1 drop; it matches Adaptive-RAG's F1 at 37.30% lower cost; and it attains up to 2.33× higher Quality-per-Monetary-Cost. Consistent gains on MuSiQue, 2WikiMultiHopQA, CRAG, and Bamboogle confirm strong out-of-distribution robustness, setting a new paradigm for finegrained resource control in multi-hop RAG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on17
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
Related papers
- CompactRAG: Reducing LLM Calls and Token Overhead in Multi-Hop Question AnsweringHao Yang, Zhiyu Yang, Xupeng Zhang, Wei Wei et al.WWW 2026
- FrugalRAG: Less is More in RL Finetuning for Multi-hop Question AnsweringAbhinav Java, Srivathsan Koundinyan, Nagarajan Natarajan, Amit SharmaICLR 2026 · 2 citations
- S2G-RAG: Structured Sufficiency and Gap Judging for Iterative Retrieval-Augmented QAMinghan Li, Junjie Zou, Xinxuan Lv, Chao Zhang et al.ACL 2026 · 1 citation
- HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented GenerationPeilin Wu, Mian Zhang, Kun Wan, Wentian Zhao et al.ICLR 2026 · 13 citations
- LiR3AG: A Lightweight Rerank Reasoning Strategy Framework for Retrieval-Augmented GenerationGuo Chen, Junjie Huang, Huaijin Xie, Fei Sun et al.AAAI 2026
