ACE: Off-Policy Actor-Critic with Causality-Aware Entropy Regularization
Tianying Ji, Yongyuan Liang, Yan Zeng, Yu Luo, Guowei Xu, Jiawei Guo, Ruijie Zheng, Furong Huang, Fuchun Sun, Huazhe Xu
Abstract
The varying significance of distinct primitive behaviors during the policy learning process has been overlooked by prior model-free RL algorithms. Leveraging this insight, we explore the causal relationship between different action dimensions and rewards to evaluate the significance of various primitive behaviors during training. We introduce a causality-aware entropy term that effectively identifies and prioritizes actions with high potential impacts for efficient exploration. Furthermore, to prevent excessive focus on specific primitive behaviors, we analyze the gradient dormancy phenomenon and introduce a dormancy-guided reset mechanism to further enhance the efficacy of our method. Our proposed algorithm, ACE: Off-policy Actor-critic with Causality-aware Entropy regularization, demonstrates a substantial performance advantage across 29 diverse continuous control tasks spanning 7 domains compared to model-free RL baselines, which underscores the effectiveness, versatility, and efficient sample efficiency of our approach. Benchmark results and videos are available at https://ace-rl.github.io/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 372bba7a-2edd-435c-999a-17ff325f72f3Cited by top-tier papers10
- Measure gradients, not activations! Enhancing neuronal activity in deep reinforcement learningJiashun Liu, Zihao Wu, Johan S. Obando-Ceron, Pablo Samuel Castro et al.NeurIPS 2025 · 15 citations
- Make-An-Agent: A Generalizable Policy Network Generator with Behavior-Prompted DiffusionYongyuan Liang, Tingqiang Xu, Kaizhe Hu, Guangqi Jiang et al.NeurIPS 2024 · 15 citations
- Premier-TACO is a Few-Shot Policy Learner: Pretraining Multitask Representation via Temporal Action-Driven Contrastive LossRuijie Zheng, Yongyuan Liang, Xiyao Wang, Shuang Ma et al.ICML 2024 · 11 citations
- BECAUSE: Bilinear Causal Representation for Generalizable Offline Model-based Reinforcement LearningHaohong Lin, Wenhao Ding, Jian Chen, Laixi Shi et al.NeurIPS 2024 · 5 citations
- Rethinking Bimanual Robotic Manipulation: Learning with Decoupled Interaction FrameworkJian-Jian Jiang, Xiao-Ming Wu, Yi-Xiang He, Ling-An Zeng et al.ICCV 2025 · 4 citations
Builds on22
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel et al.ICML 2020 · 489 citations
- Never Give Up: Learning Directed Exploration StrategiesAdrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo et al.ICLR 2020 · 349 citations
- On Warm-Starting Neural Network TrainingJordan T. Ash, Ryan P. AdamsNeurIPS 2020 · 288 citations
- The Primacy Bias in Deep Reinforcement LearningEvgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon et al.ICML 2022 · 269 citations
- Reward-Free Exploration for Reinforcement LearningChi Jin, Akshay Krishnamurthy, Max Simchowitz, Tiancheng YuICML 2020 · 226 citations
Related papers
- Promoting Stochasticity for Expressive Policies via a Simple and Efficient Regularization MethodQi Zhou, Yufei Kuang, Zherui Qiu, Houqiang Li et al.NeurIPS 2020 · 9 citations
- Striving for Simplicity and Performance in Off-Policy DRL: Output Normalization and Non-Uniform SamplingChe Wang, Yanqiu Wu, Quan Vuong, Keith W. RossICML 2020 · 38 citations
- Direct Advantage EstimationHsiao-Ru Pan, Nico Gürtler, Alexander Neitz, Bernhard SchölkopfNeurIPS 2022 · 20 citations
- Causal Information Prioritization for Efficient Reinforcement LearningHongye Cao, Fan Feng, Tianpei Yang, Jing Huo et al.ICLR 2025 · 1 citation
- Diversity Actor-Critic: Sample-Aware Entropy Regularization for Sample-Efficient ExplorationSeungyul Han, Youngchul SungICML 2021 · 34 citations
