Model-based Reinforcement Learning for Parameterized Action Spaces
Renhao Zhang, Haotian Fu, Yilin Miao, George Konidaris
摘要
We propose a novel model-based reinforcement learning algorithm -- Dynamics Learning and predictive control with Parameterized Actions (DLPA) -- for Parameterized Action Markov Decision Processes (PAMDPs). The agent learns a parameterized-action-conditioned dynamics model and plans with a modified Model Predictive Path Integral control. We theoretically quantify the difference between the generated trajectory and the optimal trajectory during planning in terms of the value they achieved through the lens of Lipschitz Continuity. Our empirical results on several standard benchmarks show that our algorithm achieves superior sample efficiency and asymptotic performance than state-of-the-art PAMDP methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Near-Optimal Second-Order Guarantees for Model-Based Adversarial Imitation LearningShangzhe Li, Dongruo Zhou, Weitong ZhangICLR 2026 · 被引用 2 次
- CHPO: Constrained Hybrid-action Policy Optimization for Reinforcement LearningAo Zhou, Jiayi Guan, Li Shen, Fan Lu 等NeurIPS 2025 · 被引用 1 次
- Learning Parameterized Skills from DemonstrationsVedant Gupta, Haotian Fu, Calvin Luo, Yiding Jiang 等NeurIPS 2025 · 被引用 1 次
- CHDP: Cooperative Hybrid Diffusion Policies for Reinforcement Learning in Parameterized Action SpaceBingyi Liu, Jinbo He, Haiyong Shi, Enshu Wang 等AAAI 2026
- From Noise to Control: Parameterized Diffusion PoliciesRenhao Zhang, Haotian Fu, Mingxi Jia, George Konidaris 等ICML 2026
它引用的顶会 Paper9
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski 等ICLR 2020 · 被引用 969 次
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel 等ICML 2020 · 被引用 489 次
相关 Paper
- Temporal Difference Learning for Model Predictive ControlNicklas Hansen, Hao Su, Xiaolong WangICML 2022 · 被引用 388 次
- Predictable MDP Abstraction for Unsupervised Model-Based RLSeohong Park, Sergey LevineICML 2023 · 被引用 11 次
- Live in the Moment: Learning Dynamics Model Adapted to Evolving PolicyXiyao Wang, Wichayaporn Wongkamjan, Ruonan Jia, Furong HuangICML 2023 · 被引用 20 次
- Rich-Observation Reinforcement Learning with Continuous Latent DynamicsYuda Song, Lili Wu, Dylan J. Foster, Akshay KrishnamurthyICML 2024 · 被引用 2 次
- Making Better Decision by Directly Planning in Continuous ControlJinhua Zhu, Yue Wang, Lijun Wu, Tao Qin 等ICLR 2023
