HyAR: Addressing Discrete-Continuous Action Reinforcement Learning via Hybrid Action Representation
Boyan Li, Hongyao Tang, Yan Zheng, Jianye Hao, Pengyi Li, Zhen Wang, Zhaopeng Meng, Li Wang
摘要
Discrete-continuous hybrid action space is a natural setting in many practical problems, such as robot control and game AI. However, most previous Reinforcement Learning (RL) works only demonstrate the success in controlling with either discrete or continuous action space, while seldom take into account the hybrid action space. One naive way to address hybrid action RL is to convert the hybrid action space into a unified homogeneous action space by discretization or continualization, so that conventional RL algorithms can be applied. However, this ignores the underlying structure of hybrid action space and also induces the scalability issue and additional approximation difficulties, thus leading to degenerated results. In this paper, we propose Hybrid Action Representation (HyAR) to learn a compact and decodable latent representation space for the original hybrid action space. HyAR constructs the latent space and embeds the dependence between discrete action and continuous parameter via an embedding table and conditional Variantional Auto-Encoder (VAE). To further improve the effectiveness, the action representation is trained to be semantically smooth through unsupervised environmental dynamics prediction. Finally, the agent then learns its policy with conventional DRL algorithms in the learned representation space and interacts with the environment by decoding the hybrid action embeddings to the original action space. We evaluate HyAR in a variety of environments with discrete-continuous action space. The results demonstrate the superiority of HyAR when compared with previous baselines, especially for high-dimensional action spaces.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- OPPerTune: Post-Deployment Configuration Tuning of Services Made EasyGagan Somashekar, Karan Tandon, Anush Kini, Chieh-Chun Chang 等NSDI 2024 · 被引用 24 次
- Improving Deep Reinforcement Learning by Reducing the Chain Effect of Value and Policy ChurnHongyao Tang, Glen BersethNeurIPS 2024 · 被引用 23 次
- OVD-Explorer: Optimism Should Not Be the Sole Pursuit of Exploration in Noisy EnvironmentsJinyi Liu, Zhi Wang, Yan Zheng, Jianye Hao 等AAAI 2024 · 被引用 14 次
- FlexPlanner: Flexible 3D Floorplanning via Deep Reinforcement Learning in Hybrid Action Space with Multi-Modality RepresentationRuizhe Zhong, Xingbo Du, Shixiong Kai, Zhentao Tang 等NeurIPS 2024 · 被引用 8 次
- Generalized Policy Iteration using Tensor Approximation for Hybrid ControlSuhan Shetty, Teng Xue, Sylvain CalinonICLR 2024 · 被引用 8 次
它引用的顶会 Paper7
- Decoupling Representation Learning from Reinforcement LearningAdam Stooke, Kimin Lee, Pieter Abbeel, Michael LaskinICML 2021 · 被引用 389 次
- Reinforcement Learning with Prototypical RepresentationsDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICML 2021 · 被引用 262 次
- Improving black-box optimization in VAE latent space using decoder uncertaintyPascal Notin, José Miguel Hernández-Lobato, Yarin GalNeurIPS 2021 · 被引用 76 次
- Automatic Web Testing Using Curiosity-Driven Reinforcement LearningYan Zheng, Yi Liu, Xiaofei Xie, Yepang Liu 等ICSE 2021 · 被引用 75 次
- Continuous Multiagent Control Using Collective Behavior Entropy for Large-Scale Home Energy ManagementJianwen Sun, Yan Zheng, Jianye Hao, Zhaopeng Meng 等AAAI 2020 · 被引用 20 次
相关 Paper
- Efficient Planning in a Compact Latent Action SpaceZhengyao Jiang, Tianjun Zhang, Michael Janner, Yueying Li 等ICLR 2023 · 被引用 3 次
- CHDP: Cooperative Hybrid Diffusion Policies for Reinforcement Learning in Parameterized Action SpaceBingyi Liu, Jinbo He, Haiyong Shi, Enshu Wang 等AAAI 2026
- Learning High-Frequency Continuous Action Chunks in Latent SpaceKunyun Wang, Yuhang Zheng, Yupeng Zheng, Jieru Zhao 等ICML 2026
- Distributions as Actions: A Unified Framework for Diverse Action SpacesJiamin He, A. Rupam Mahmood, Martha WhiteICLR 2026
- Identifying latent state transitions in non-linear dynamical systemsÇaglar Hizli, Çagatay Yildiz, Matthias Bethge, S. T. John 等ICLR 2025
