Controlling Underestimation Bias in Reinforcement Learning via Quasi-median Operation
Wei Wei, Yujia Zhang, Jiye Liang, Lin Li, Yuze Li
Abstract
How to get a good value estimation is one of the key problems in reinforcement learning (RL). Current off-policy methods, such as Maxmin Q-learning, TD3, and TADD, suffer from the underestimation problem when solving the overestimation problem. In this paper, we propose the Quasi-Median Operation, a novel way to mitigate the underestimation bias by selecting the quasi-median from multiple state-action values. Based on the quasi-median operation, we propose Quasi-Median Q-learning (QMQ) for the discrete action tasks and Quasi-Median Delayed Deep Deterministic Policy Gradient (QMD3) for the continuous action tasks. Theoretically, the underestimation bias of our method is improved while the estimation variance is significantly reduced compared to Maxmin Q-learning, TD3, and TADD. We conduct extensive experiments on the discrete and continuous action tasks, and results show that our method outperforms the state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ab63361-d125-40b2-b5c6-a4b29b5b2624Cited by top-tier papers4
- Keep Various Trajectories: Promoting Exploration of Ensemble Policies in Continuous ControlChao Li, Chen Gong, Qiang He, Xinwen HouNeurIPS 2023 · 8 citations
- Offline-Boosted Actor-Critic: Adaptively Blending Optimal Historical Behaviors in Deep Off-Policy RLYu Luo, Tianying Ji, Fuchun Sun, Jianwei Zhang et al.ICML 2024 · 7 citations
- Dynamic Uncertainty Estimation for Offline Reinforcement LearningJiesheng Wang, Lin Li, Wei Wei, Yujia Zhang et al.AAAI 2025 · 2 citations
- Efficient Offline Reinforcement Learning via Peer-Influenced ConstraintYujia Zhang, Lin Li, Wei Wei, Jianguo Wu et al.ICLR 2026
Builds on3
- Maxmin Q-learning: Controlling the Estimation Bias of Q-learningQingfeng Lan, Yangchen Pan, Alona Fyshe, Martha WhiteICLR 2020 · 213 citations
- Hierarchical Reinforcement Learning for Integrated RecommendationRuobing Xie, Shaoliang Zhang, Rui Wang, Feng Xia et al.AAAI 2021 · 90 citations
- A General Offline Reinforcement Learning Framework for Interactive RecommendationTeng Xiao, Donglin WangAAAI 2021 · 82 citations
Related papers
- Action Candidate Based Clipped Double Q-learning for Discrete and Continuous Action TasksHaobo Jiang, Jin Xie, Jian YangAAAI 2021 · 20 citations
- Softmax Deep Double Deterministic Policy GradientsLing Pan, Qingpeng Cai, Longbo HuangNeurIPS 2020 · 138 citations
- ADDQ: Adaptive distributional double Q-learningLeif Döring, Benedikt Wille, Maximilian Birr, Mihail Bîrsan et al.ICML 2025
- Efficient Continuous Control with Double Actors and Regularized CriticsJiafei Lyu, Xiaoteng Ma, Jiangpeng Yan, Xiu LiAAAI 2022 · 69 citations
- Self-correcting Q-learningRong Zhu, Mattia RigottiAAAI 2021 · 22 citations
