Reinforcement Learning of Causal Variables Using Mediation Analysis
Tue Herlau, Rasmus Larsen
Abstract
We consider the problem of acquiring causal representations and concepts in a reinforcement learning setting. Our approach defines a causal variable as being both manipulable by a policy, and able to predict the outcome. We thereby obtain a parsimonious causal graph in which interventions occur at the level of policies. The approach avoids defining a generative model of the data, prior pre-processing, or learning the transition kernel of the Markov decision process. Instead, causal variables and policies are determined by maximizing a new optimization target inspired by mediation analysis, which differs from the expected return. The maximization is accomplished using a generalization of Bellman's equation which is shown to converge, and the method finds meaningful causal representations in a simulated environment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a6bfff0e-2f61-44cd-a575-58005cd1d6c9Cited by top-tier papers1
Ask how each one uses itBuilds on2
- Counterfactuals uncover the modular structure of deep generative modelsMichel Besserve, Arash Mehrjou, Rémy Sun, Bernhard SchölkopfICLR 2020 · 109 citations
- Provably Efficient Causal Reinforcement Learning with Confounded Observational DataLingxiao Wang, Zhuoran Yang, Zhaoran WangNeurIPS 2021 · 61 citations
Related papers
- Generalizing Goal-Conditioned Reinforcement Learning with Variational Causal ReasoningWenhao Ding, Haohong Lin, Bo Li, Ding ZhaoNeurIPS 2022 · 59 citations
- Hierarchical Reinforcement Learning with Targeted Causal InterventionsMohammadsadegh Khorasani, Saber Salehkaleybar, Negar Kiyavash, Matthias GrossglauserICML 2025
- Reward-oriented Causal Representation LearningZirui Yan, Emre Acartürk, Ali TajerNeurIPS 2025
- Learning to Perceive the World Through Control: Empowerment-Based Representation LearningMahsa Bastankhah, Sophie Broderick, Benjamin EysenbachICML 2026
- Contextual Causal Bayesian OptimisationVahan Arsenyan, Antoine Grosnit, Haitham Bou-Ammar, Arnak S. DalalyanICLR 2026 · 3 citations
