The Statistical Benefits of Quantile Temporal-Difference Learning for Value Estimation
Mark Rowland, Yunhao Tang, Clare Lyle, Rémi Munos, Marc G. Bellemare, Will Dabney
2023年份
13被引次数
12顶会引用
摘要
We study the problem of temporal-difference-based policy evaluation in reinforcement learning. In particular, we analyse the use of a distributional reinforcement learning algorithm, quantile temporal-difference learning (QTD), for this task. We reach the surprising conclusion that even if a practitioner has no interest in the return distribution beyond the mean, QTD (which learns predictions about the full distribution of returns) may offer performance superior to approaches such as classical TD learning, which predict only the mean return, even in the tabular setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Stop Regressing: Training Value Functions via Classification for Scalable Deep RLJesse Farebrother, Jordi Orbay, Quan Vuong, Adrien Ali Taïga 等ICML 2024 · 被引用 118 次
- More Benefits of Being Distributional: Second-Order Bounds for Reinforcement LearningKaiwen Wang, Owen Oertell, Alekh Agarwal, Nathan Kallus 等ICML 2024 · 被引用 20 次
- Q#: Provably Optimal Distributional RL for LLM Post-TrainingJin Peng Zhou, Kaiwen Wang, Jonathan D. Chang, Zhaolin Gao 等NeurIPS 2025 · 被引用 18 次
- Beyond Average Return in Markov Decision ProcessesAlexandre Marthe, Aurélien Garivier, Claire VernadeNeurIPS 2023 · 被引用 15 次
- A Finite Sample Analysis of Distributional TD Learning with Linear Function ApproximationYang Peng, Kaicheng Jin, Liangyu Zhang, Zhihua ZhangNeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper9
- Munchausen Reinforcement LearningNino Vieillard, Olivier Pietquin, Matthieu GeistNeurIPS 2020 · 被引用 120 次
- The Value-Improvement Path: Towards Better Representations for Reinforcement LearningWill Dabney, André Barreto, Mark Rowland, Robert Dadashi 等AAAI 2021 · 被引用 76 次
- Muesli: Combining Improvements in Policy OptimizationMatteo Hessel, Ivo Danihelka, Fabio Viola, Arthur Guez 等ICML 2021 · 被引用 69 次
- Non-Crossing Quantile Regression for Distributional Reinforcement LearningFan Zhou, Jianing Wang, Xingdong FengNeurIPS 2020 · 被引用 63 次
- Gradient Temporal-Difference Learning with Regularized CorrectionsSina Ghiassian, Andrew Patterson, Shivam Garg, Dhawal Gupta 等ICML 2020 · 被引用 49 次
相关 Paper
- A Differential Perspective on Distributional Reinforcement LearningJuan Sebastian Rojas, Chi-Guhn LeeAAAI 2026 · 被引用 4 次
- A Principled Path to Fitted Distributional EvaluationSungee Hong, Jiayi Wang, Zhengling Qi, Raymond K. W. WongNeurIPS 2025
- Distributional Bellman Operators over Mean EmbeddingsLi Kevin Wenliang, Grégoire Delétang, Matthew Aitchison, Marcus Hutter 等ICML 2024 · 被引用 5 次
- Distributional Reinforcement Learning with Monotonic SplinesYudong Luo, Guiliang Liu, Haonan Duan, Oliver Schulte 等ICLR 2022 · 被引用 18 次
- Statistical Efficiency of Distributional Temporal Difference LearningYang Peng, Liangyu Zhang, Zhihua ZhangNeurIPS 2024 · 被引用 8 次
