Turning Sand to Gold: Recycling Data to Bridge On-Policy and Off-Policy Learning via Causal Bound
Tal Fiskus, Uri Shaham
Abstract
Deep reinforcement learning (DRL) agents excel in solving complex decisionmaking tasks across various domains. However, they often require a substantial number of training steps and a vast experience replay buffer, leading to significant computational and resource demands. To address these challenges, we introduce a novel theoretical result that leverages the Neyman-Rubin potential outcomes framework into DRL. Unlike most methods that focus on bounding the counterfactual loss, we establish a causal bound on the factual loss, which is analogous to the on-policy loss in DRL. This bound is computed by storing past value network outputs in the experience replay buffer, effectively utilizing data that is usually discarded. Extensive experiments across the Atari 2600 and MuJoCo domains on various agents, such as DQN and SAC, achieve up to 383% higher reward ratio, outperforming the same agents without our proposed term, and reducing the experience replay buffer size by up to 96%, significantly improving sample efficiency at a negligible cost 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on20
- Agent57: Outperforming the Atari Human BenchmarkAdrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann et al.ICML 2020 · 584 citations
- Mastering Atari Games with Limited DataWeirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel et al.NeurIPS 2021 · 345 citations
- Causal Transformer for Estimating Counterfactual OutcomesValentyn Melnychuk, Dennis Frauen, Stefan FeuerriegelICML 2022 · 146 citations
- Causal Influence Detection for Improving Efficiency in Reinforcement LearningMaximilian Seitzer, Bernhard Schölkopf, Georg MartiusNeurIPS 2021 · 120 citations
- Off-policy Policy Evaluation For Sequential Decisions Under Unobserved ConfoundingHongseok Namkoong, Ramtin Keramati, Steve Yadlowsky, Emma BrunskillNeurIPS 2020 · 81 citations
Related papers
- How Does Goal Relabeling Improve Sample Efficiency?Sirui Zheng, Chenjia Bai, Zhuoran Yang, Zhaoran WangICML 2024 · 5 citations
- The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement LearningJiashun Liu, Johan S. Obando-Ceron, Pablo Samuel Castro, Aaron C. Courville et al.ICML 2025
- Neural Episodic Control with State AbstractionZhuo Li, Derui Zhu, Yujing Hu, Xiaofei Xie et al.ICLR 2023 · 5 citations
- Large Batch Experience ReplayThibault Lahire, Matthieu Geist, Emmanuel RachelsonICML 2022 · 18 citations
- [CASPI] Causal-aware Safe Policy Improvement for Task-oriented DialogueGovardana Sachithanandam Ramachandran, Kazuma Hashimoto, Caiming XiongACL 2022 · 12 citations
