Towards Empowerment Gain through Causal Structure Learning in Model-Based Reinforcement Learning
Hongye Cao, Fan Feng, Meng Fang, Shaokang Dong, Tianpei Yang, Jing Huo, Yang Gao
Abstract
In Model-Based Reinforcement Learning (MBRL), incorporating causal structures into dynamics models provides agents with the structured understanding of environments, enabling more efficient and effective decisions. Empowerment, as an intrinsic motivation, enhances the ability of agents to actively control environments by maximizing mutual information between future states and actions. We posit that empowerment coupled with the causal understanding of the environment can improve the agent's controllability over environments, while enhanced empowerment gain can further facilitate causal reasoning. To this end, we propose the framework that pioneers the integration of empowerment with causal reasoning, Empowerment through Causal Learning (ECL), where an agent with the awareness of the causal dynamics model achieves empowerment-driven exploration and optimizes its causal structure for task learning. Specifically, we first train a causal dynamics model of the environment based on collected data. Next, we maximize empowerment under the causal structure for exploration, simultaneously using data gathered through exploration to update the causal dynamics model, which could be more controllable than dynamics models without the causal structure. We also design an intrinsic curiosity reward to mitigate overfitting during downstream task learning. Importantly, ECL is method-agnostic and can integrate diverse causal discovery methods. We evaluate ECL combined with 3 causal discovery methods across 6 environments including both state-based and pixel-based tasks, demonstrating its performance gain compared to other causal MBRL methods, in terms of causal structure discovery, sample efficiency, and asymptotic performance in policy learning. The project page is https://sites.google.com/view/ecl-1429/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak AttacksHongye Cao, Sijia Jing, Yanming Wang, Ziyue Peng et al.ICLR 2026 · 28 citations
- Task-Aware Exploration via a Predictive Bisimulation MetricDayang Liang, Ruihan LIU, Lipeng Wan, Yunlong Liu et al.ICML 2026 · 1 citation
- Curious Causality-Seeking Agents in Open-ended WorldsZhiyu Zhao, Haoxuan Li, Haifeng Zhang, Jun Wang et al.NeurIPS 2025
Builds on23
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- Causal Influence Detection for Improving Efficiency in Reinforcement LearningMaximilian Seitzer, Bernhard Schölkopf, Georg MartiusNeurIPS 2021 · 120 citations
- Invariant Causal Representation Learning for Out-of-Distribution GeneralizationChaochao Lu, Yuhuai Wu, José Miguel Hernández-Lobato, Bernhard SchölkopfICLR 2022 · 119 citations
- Robust agents learn causal world modelsJonathan Richens, Tom EverittICLR 2024 · 78 citations
- Learning Invariant Representations for Reinforcement Learning without ReconstructionAmy Zhang, Rowan Thomas McAllister, Roberto Calandra, Yarin Gal et al.ICLR 2021 · 77 citations
Related papers
- MetaCURE: Meta Reinforcement Learning with Empowerment-Driven ExplorationJin Zhang, Jianhao Wang, Hao Hu, Tong Chen et al.ICML 2021 · 33 citations
- Learning to Perceive the World Through Control: Empowerment-Based Representation LearningMahsa Bastankhah, Sophie Broderick, Benjamin EysenbachICML 2026
- Causal Information Prioritization for Efficient Reinforcement LearningHongye Cao, Fan Feng, Tianpei Yang, Jing Huo et al.ICLR 2025 · 1 citation
- Causality-Aware Efficient Exploration for Cooperative Multi-Agent Reinforcement LearningHongye Cao, Tianpei Yang, Fan Feng, Hammadi Rafik Ouariachi et al.AAAI 2026
- Information Prioritization through Empowerment in Visual Model-based RLHomanga Bharadhwaj, Mohammad Babaeizadeh, Dumitru Erhan, Sergey LevineICLR 2022 · 35 citations
