Learning Nonlinear Causal Reductions to Explain Reinforcement Learning Policies
Armin Kekić, Jan Schneider, Dieter Büchler, Bernhard Schölkopf, Michel Besserve
摘要
Why do reinforcement learning (RL) policies fail or succeed? This is a challenging question due to the complex, high-dimensional nature of agent-environment interactions. In this work, we take a causal perspective on explaining the behavior of RL policies by viewing the states, actions, and rewards as variables in a low-level causal model. We introduce random perturbations to policy actions during execution and observe their effects on the cumulative reward, learning a simplified high-level causal model that explains these relationships. To this end, we develop a nonlinear Causal Model Reduction framework that ensures approximate interventional consistency, meaning the simplified high-level model responds to interventions in a similar way as the original complex system. We prove that for a class of nonlinear causal models, there exists a unique solution that achieves exact interventional consistency, ensuring learned explanations reflect meaningful causal patterns. Experiments on both synthetic causal models and practical RL tasks -including pendulum control and robot table tennis -demonstrate that our approach can uncover important behavioral patterns, biases, and failure modes in trained RL policies. * Joint supervision. Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Causal Abstractions of Neural NetworksAtticus Geiger, Hanson Lu, Thomas Icard, Christopher PottsNeurIPS 2021 · 被引用 516 次
- Explainable Reinforcement Learning through a Causal LensPrashan Madumal, Tim Miller, Liz Sonenberg, Frank VetereAAAI 2020 · 被引用 408 次
- Homomorphism AutoEncoder - Learning Group Structured Representations from Observed TransitionsHamza Keurti, Hsiao-Ru Pan, Michel Besserve, Benjamin F. Grewe 等ICML 2023 · 被引用 22 次
相关 Paper
- Reinforcement Learning of Causal Variables Using Mediation AnalysisTue Herlau, Rasmus LarsenAAAI 2022 · 被引用 8 次
- CausalXRL: Explainable Reinforcement Learning through Causal Graph ReasoningYanming Zhang, Eric Papenhausen, Klaus MuellerICML 2026
- Tell me why! Explanations support learning relational and causal structureAndrew K. Lampinen, Nicholas A. Roy, Ishita Dasgupta, Stephanie C. Y. Chan 等ICML 2022 · 被引用 51 次
- Interventionally Consistent Surrogates for Complex Simulation ModelsJoel Dyer, Nicholas Bishop, Yorgos Felekis, Fabio Massimo Zennaro 等NeurIPS 2024 · 被引用 12 次
- Deep Bayesian Nonparametric Learning of Rules and Plans from Demonstrations with a Learned Automaton PriorBrandon Araki, Kiran Vodrahalli, Thomas Leech, Cristian Ioan Vasile 等AAAI 2020 · 被引用 8 次
