Action Candidate Based Clipped Double Q-learning for Discrete and Continuous Action Tasks
Haobo Jiang, Jin Xie, Jian Yang
摘要
Double Q-learning is a popular reinforcement learning algorithm in Markov decision process (MDP) problems. Clipped Double Q-learning, as an effective variant of Double Qlearning, employs the clipped double estimator to approximate the maximum expected action value. Due to the underestimation bias of the clipped double estimator, performance of clipped Double Q-learning may be degraded in some stochastic environments. In this paper, in order to reduce the underestimation bias, we propose an action candidate based clipped double estimator for Double Q-learning. Specifically, we first select a set of elite action candidates with the high action values from one set of estimators. Then, among these candidates, we choose the highest valued action from the other set of estimators. Finally, we use the maximum value in the second set of estimators to clip the action value of the chosen action in the first set of estimators and the clipped value is used for approximating the maximum expected action value. Theoretically, the underestimation bias in our clipped Double Q-learning decays monotonically as the number of the action candidates decreases. Moreover, the number of action candidates controls the trade-off between the overestimation and underestimation biases. In addition, we also extend our clipped Double Q-learning to continuous action tasks via approximating the elite continuous action candidates. We empirically verify that our algorithm can more accurately estimate the maximum expected action value on some toy environments and yield good performance on several benchmark problems. All code and hyperparameters available at https://github.com/Jiang-HB/AC CDQ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM ReasoningKongcheng Zhang, Qi Yao, Shunyu Liu, Yingjie Wang 等NeurIPS 2025 · 被引用 45 次
- NaviFormer: A Spatio-Temporal Context-Aware Transformer for Object NavigationWei Xie, Haobo Jiang, Yun Zhu, Jianjun Qian 等AAAI 2025 · 被引用 7 次
- A Simple "Motivation" Can Enhance Reinforcement Finetuning of Large Reasoning ModelsJunjie Zhang, Guozheng Ma, Shunyu Liu, Haoyu Wang 等ICLR 2026 · 被引用 6 次
- FUSER: Feed-Forward Multiview 3D Registration Transformer and SE(3)^N Diffusion RefinementHaobo Jiang, Jin Xie, Jian Yang, Liang Yu 等CVPR 2026 · 被引用 5 次
- ReACT: Reward-informed Autoregressive Decision CAD TransformerYijie Ding, Yang Liu, Haobo Jiang, Jianmin ZhengAAAI 2026 · 被引用 1 次
相关 Paper
- Ensemble Bootstrapping for Q-LearningOren Peer, Chen Tessler, Nadav Merlis, Ron MeirICML 2021 · 被引用 56 次
- Self-correcting Q-learningRong Zhu, Mattia RigottiAAAI 2021 · 被引用 22 次
- On the Estimation Bias in Double Q-LearningZhizhou Ren, Guangxiang Zhu, Hao Hu, Beining Han 等NeurIPS 2021 · 被引用 35 次
- Controlling Underestimation Bias in Reinforcement Learning via Quasi-median OperationWei Wei, Yujia Zhang, Jiye Liang, Lin Li 等AAAI 2022 · 被引用 20 次
- Maxmin Q-learning: Controlling the Estimation Bias of Q-learningQingfeng Lan, Yangchen Pan, Alona Fyshe, Martha WhiteICLR 2020 · 被引用 213 次
