Non-Crossing Quantile Regression for Distributional Reinforcement Learning
Fan Zhou, Jianing Wang, Xingdong Feng
Abstract
Distributional reinforcement learning (DRL) estimates the distribution over future returns instead of the mean to more efficiently capture the intrinsic uncertainty of MDPs. However, batch-based DRL algorithms cannot guarantee the non-decreasing property of learned quantile curves especially at the early training stage, leading to abnormal distribution estimates and reduced model interpretability. To address these issues, we introduce a general DRL framework by using non-crossing quantile regression to ensure the monotonicity constraint within each sampled batch, which can be incorporated with some well-known DRL algorithm. We demonstrate the validity of our method from both the theory and model implementation perspectives. Experiments on Atari 2600 Games show that some state-of-art DRL algorithms with the non-crossing modification can significantly outperform their baselines in terms of faster convergence speeds and better testing performance. In particular, our method can effectively recover the distribution information and thus dramatically increase the exploration efficiency when the reward space is extremely sparse.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 19607479-fb39-485e-8832-a5e9d81d3f5fCited by top-tier papers14
- RiskQ: Risk-sensitive Multi-Agent Reinforcement Learning Value FactorizationSiqi Shen, Chennan Ma, Chao Li, Weiquan Liu et al.NeurIPS 2023 · 34 citations
- The Sufficiency of Off-Policyness and Soft Clipping: PPO Is Still Insufficient according to an Off-Policy MeasureXing Chen, Dongcui Diao, Hechang Chen, Hengshuai Yao et al.AAAI 2023 · 28 citations
- PI3NN: Out-of-distribution-aware Prediction Intervals from Three Neural NetworksSiyan Liu, Pei Zhang, Dan Lu, Guannan ZhangICLR 2022 · 13 citations
- The Statistical Benefits of Quantile Temporal-Difference Learning for Value EstimationMark Rowland, Yunhao Tang, Clare Lyle, Rémi Munos et al.ICML 2023 · 13 citations
- Uncertainty-Aware Reinforcement Learning for Risk-Sensitive Player Evaluation in Sports GameGuiliang Liu, Yudong Luo, Oliver Schulte, Pascal PoupartNeurIPS 2022 · 10 citations
Related papers
- Distributional Reinforcement Learning with Monotonic SplinesYudong Luo, Guiliang Liu, Haonan Duan, Oliver Schulte et al.ICLR 2022 · 18 citations
- Variance Control for Distributional Reinforcement LearningQi Kuang, Zhoufan Zhu, Liwen Zhang, Fan ZhouICML 2023 · 4 citations
- A Differential Perspective on Distributional Reinforcement LearningJuan Sebastian Rojas, Chi-Guhn LeeAAAI 2026 · 4 citations
- Bayesian Distributional Policy GradientsLuchen Li, A. Aldo FaisalAAAI 2021 · 11 citations
- The Nature of Temporal Difference Errors in Multi-step Distributional Reinforcement LearningYunhao Tang, Rémi Munos, Mark Rowland, Bernardo Ávila Pires et al.NeurIPS 2022 · 16 citations
