No Prior Mask: Eliminate Redundant Action for Deep Reinforcement Learning
Dianyu Zhong, Yiqin Yang, Qianchuan Zhao
摘要
The large action space is one fundamental obstacle to deploying Reinforcement Learning methods in the real world. The numerous redundant actions will cause the agents to make repeated or invalid attempts, even leading to task failure. Although current algorithms conduct some initial explorations for this issue, they either suffer from rule-based systems or depend on expert demonstrations, which significantly limits their applicability in many real-world settings. In this work, we examine the theoretical analysis of what action can be eliminated in policy optimization and propose a novel redundant action filtering mechanism. Unlike other works, our method constructs the similarity factor by estimating the distance between the state distributions, which requires no prior knowledge. In addition, we combine the modified inverse model to avoid extensive computation in high-dimensional state space. We reveal the underlying structure of action spaces and propose a simple yet efficient redundant action filtering mechanism named No Prior Mask (NPM) based on the above techniques. We show the superior performance of our method by conducting extensive experiments on high-dimensional, pixel-input, and stochastic problems with various action redundancy tasks. Our code is public online at https://github.com/zhongdy15/npm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- On Entropy Control in LLM-RL AlgorithmsHan ShenICLR 2026 · 被引用 43 次
- Excluding the Irrelevant: Focusing Reinforcement Learning through Continuous Action MaskingRoland Stolz, Hanna Krasowski, Jakob Thumm, Michael Eichelbeck 等NeurIPS 2024 · 被引用 30 次
- xMTF: A Formula-Free Model for Reinforcement-Learning-Based Multi-Task Fusion in Recommender SystemsYang Cao, Changhao Zhang, Xiaoshuang Chen, Kaiqiao Zhan 等WWW 2025 · 被引用 3 次
- STAIR: Addressing Stage Misalignment through Temporal-Aligned Preference Reinforcement LearningYao Luan, Ni Mu, Yiqin Yang, Bo Xu 等NeurIPS 2025 · 被引用 3 次
- Enhancing Diversity In Parallel Agents: A Maximum State Entropy Exploration StoryVincenzo De Paola, Riccardo Zamboni, Mirco Mutti, Marcello RestelliICML 2025
它引用的顶会 Paper7
- Mastering Complex Control in MOBA Games with Deep Reinforcement LearningDeheng Ye, Zhao Liu, Mingfei Sun, Bei Shi 等AAAI 2020 · 被引用 395 次
- Believe What You See: Implicit Constraint Approach for Offline Multi-Agent Reinforcement LearningYiqin Yang, Xiaoteng Ma, Chenghao Li, Zewu Zheng 等NeurIPS 2021 · 被引用 133 次
- Offline Reinforcement Learning with Value-based Episodic MemoryXiaoteng Ma, Yiqin Yang, Hao Hu, Jun Yang 等ICLR 2022 · 被引用 51 次
- Continuous Control with Action Quantization from DemonstrationsRobert Dadashi, Léonard Hussenot, Damien Vincent, Sertan Girgin 等ICML 2022 · 被引用 32 次
- Learning to Represent Action Values as a Hypergraph on the Action VerticesArash Tavakoli, Mehdi Fatemi, Petar KormushevICLR 2021 · 被引用 25 次
相关 Paper
- Achieving Sample and Computational Efficient Reinforcement Learning by Action Space Reduction via GroupingYining Li, Peizhong Ju, Ness B. ShroffICLR 2024 · 被引用 1 次
- Provably Filtering Exogenous Distractors using Multistep Inverse DynamicsYonathan Efroni, Dipendra Misra, Akshay Krishnamurthy, Alekh Agarwal 等ICLR 2022 · 被引用 38 次
- Learning Pseudometric-based Action Representations for Offline Reinforcement LearningPengjie Gu, Mengchen Zhao, Chen Chen, Dong Li 等ICML 2022 · 被引用 17 次
- Language Model Adaption for Reinforcement Learning with Natural Language Action SpaceJiangxing Wang, Jiachen Li, Xiao Han, Deheng Ye 等ACL 2024
- CSO: Constraint-Guided Space Optimization for Active Scene MappingXuefeng Yin, Chenyang Zhu, Shanglai Qu, Yuqi Li 等ACM MM 2024
