The surprising efficiency of temporal difference learning for rare event prediction
Xiaoou Cheng, Jonathan Weare
摘要
We quantify the efficiency of temporal difference (TD) learning over the direct, or Monte Carlo (MC), estimator for policy evaluation in reinforcement learning, with an emphasis on estimation of quantities related to rare events. Policy evaluation is complicated in the rare event setting by the long timescale of the event and by the need for relative accuracy in estimates of very small values. Specifically, we focus on least-squares TD (LSTD) prediction for finite state Markov chains, and show that LSTD can achieve relative accuracy far more efficiently than MC. We prove a central limit theorem for the LSTD estimator and upper bound the relative asymptotic variance by simple quantities characterizing the connectivity of states relative to the transition probabilities between them. Using this bound, we show that, even when both the timescale of the rare event and the relative accuracy of the MC estimator are exponentially large in the number of states, LSTD maintains a fixed level of relative accuracy with a total number of observed transitions of the Markov chain that is only polynomially large in the number of states.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper1
相关 Paper
- Uncorrected Least-Squares Temporal Difference with Lambda-ReturnTakayuki OsogamiAAAI 2020 · 被引用 1 次
- Managing Temporal Resolution in Continuous Value Estimation: A Fundamental Trade-offZichen Vincent Zhang, Johannes Kirschner, Junxi Zhang, Francesco Zanini 等NeurIPS 2023 · 被引用 3 次
- The Statistical Benefits of Quantile Temporal-Difference Learning for Value EstimationMark Rowland, Yunhao Tang, Clare Lyle, Rémi Munos 等ICML 2023 · 被引用 13 次
- Variance Reduction via Resampling and Experience ReplayJiale Han, Xiaowu Dai, Yuhua ZhuAAAI 2026
- Reanalysis of Variance Reduced Temporal Difference LearningTengyu Xu, Zhe Wang, Yi Zhou, Yingbin LiangICLR 2020 · 被引用 46 次
