Variance Reduction via Resampling and Experience Replay
Jiale Han, Xiaowu Dai, Yuhua Zhu
Abstract
Experience replay is a foundational technique in reinforcement learning that enhances learning stability by storing past experiences in a replay buffer and reusing them during training. Despite its practical success, its theoretical properties remain underexplored. In this paper, we present a theoretical framework that models experience replay using resampled Uand V -statistics, providing rigorous variance reduction guarantees. We apply this framework to policy evaluation tasks using the Least-Squares Temporal Difference (LSTD) algorithm and a Partial Differential Equation (PDE)-based modelfree algorithm, demonstrating significant improvements in stability and efficiency, particularly in data-scarce scenarios. Beyond policy evaluation, we extend the framework to kernel ridge regression, showing that the experience replay-based method reduces the computational cost from the traditional O(n 3 ) in time to as low as O(n 2 ) in time while simultaneously reducing variance. Extensive numerical experiments validate our theoretical findings, demonstrating the broad applicability and effectiveness of experience replay in diverse machine learning tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5443ed06-75d0-4211-a4bf-3642bf37b684Builds on1
Related papers
- Large Batch Experience ReplayThibault Lahire, Matthieu Geist, Emmanuel RachelsonICML 2022 · 18 citations
- The surprising efficiency of temporal difference learning for rare event predictionXiaoou Cheng, Jonathan WeareNeurIPS 2024 · 10 citations
- Uncorrected Least-Squares Temporal Difference with Lambda-ReturnTakayuki OsogamiAAAI 2020 · 1 citation
- Learning Expected Emphatic Traces for Deep RLRay Jiang, Shangtong Zhang, Veronica Chelu, Adam White et al.AAAI 2022 · 14 citations
- Taylor TD-learningMichele Garibbo, Maxime Robeyns, Laurence AitchisonNeurIPS 2023
