Variance Reduction via Resampling and Experience Replay
Jiale Han, Xiaowu Dai, Yuhua Zhu
摘要
Experience replay is a foundational technique in reinforcement learning that enhances learning stability by storing past experiences in a replay buffer and reusing them during training. Despite its practical success, its theoretical properties remain underexplored. In this paper, we present a theoretical framework that models experience replay using resampled Uand V -statistics, providing rigorous variance reduction guarantees. We apply this framework to policy evaluation tasks using the Least-Squares Temporal Difference (LSTD) algorithm and a Partial Differential Equation (PDE)-based modelfree algorithm, demonstrating significant improvements in stability and efficiency, particularly in data-scarce scenarios. Beyond policy evaluation, we extend the framework to kernel ridge regression, showing that the experience replay-based method reduces the computational cost from the traditional O(n 3 ) in time to as low as O(n 2 ) in time while simultaneously reducing variance. Extensive numerical experiments validate our theoretical findings, demonstrating the broad applicability and effectiveness of experience replay in diverse machine learning tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- Large Batch Experience ReplayThibault Lahire, Matthieu Geist, Emmanuel RachelsonICML 2022 · 被引用 18 次
- The surprising efficiency of temporal difference learning for rare event predictionXiaoou Cheng, Jonathan WeareNeurIPS 2024 · 被引用 10 次
- Uncorrected Least-Squares Temporal Difference with Lambda-ReturnTakayuki OsogamiAAAI 2020 · 被引用 1 次
- Learning Expected Emphatic Traces for Deep RLRay Jiang, Shangtong Zhang, Veronica Chelu, Adam White 等AAAI 2022 · 被引用 14 次
- Taylor TD-learningMichele Garibbo, Maxime Robeyns, Laurence AitchisonNeurIPS 2023
