Causal Discovery with Reinforcement Learning
Shengyu Zhu, Ignavier Ng, Zhitang Chen
Abstract
Discovering causal structure among a set of variables is a fundamental problem in many empirical sciences. Traditional score-based casual discovery methods rely on various local heuristics to search for a Directed Acyclic Graph (DAG) according to a predefined score function. While these methods, e.g., greedy equivalence search, may have attractive results with infinite samples and certain model assumptions, they are usually less satisfactory in practice due to finite data and possible violation of assumptions. Motivated by recent advances in neural combinatorial optimization, we propose to use Reinforcement Learning (RL) to search for the DAG with the best scoring. Our encoder-decoder model takes observable data as input and generates graph adjacency matrices that are used to compute rewards. The reward incorporates both the predefined score function and two penalty terms for enforcing acyclicity. In contrast with typical RL applications where the goal is to learn a policy, we use RL as a search strategy and our final output would be the graph, among all graphs generated during training, that achieves the best reward. We conduct experiments on both synthetic and real datasets, and show that the proposed approach not only has an improved search ability but also allows a flexible score function under the acyclicity constraint.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6cb4ca85-a6ca-44c9-977b-6ae3c2c2cdb3Cited by top-tier papers67
- On the Role of Sparsity and DAG Constraints for Learning Linear DAGsIgnavier Ng, AmirEmad Ghassami, Kun ZhangNeurIPS 2020 · 306 citations
- Differentiable Causal Discovery from Interventional DataPhilippe Brouillard, Sébastien Lachapelle, Alexandre Lacoste, Simon Lacoste-Julien et al.NeurIPS 2020 · 295 citations
- Score Matching Enables Causal Discovery of Nonlinear Additive Noise ModelsPaul Rolland, Volkan Cevher, Matthäus Kleindessner, Chris Russell et al.ICML 2022 · 123 citations
- Amortized Inference for Causal Structure LearningLars Lorch, Scott Sussex, Jonas Rothfuss, Andreas Krause et al.NeurIPS 2022 · 118 citations
- Efficient Neural Causal Discovery without Acyclicity ConstraintsPhillip Lippe, Taco Cohen, Efstratios GavvesICLR 2022 · 95 citations
Builds on1
Related papers
- Reinforcement Causal Structure Learning on Order GraphDezhi Yang, Guoxian Yu, Jun Wang, Zhengtian Wu et al.AAAI 2023 · 20 citations
- Score-based Greedy Search for Structure Identification of Partially Observed Causal ModelsXinshuai Dong, Ignavier Ng, Haoyue Dai, Jiaqi Sun et al.ICLR 2026 · 1 citation
- MARLIN: Multi-Agent Reinforcement Learning for Incremental DAG DiscoveryDong Li, Zhengzhang Chen, Xujiang Zhao, Linlin Yu et al.AAAI 2026
- Less Greedy Equivalence SearchAdiba Ejaz, Elias BareinboimNeurIPS 2025 · 1 citation
- Integer Programming for Causal Structure Learning in the Presence of Latent VariablesRui Chen, Sanjeeb Dash, Tian GaoICML 2021 · 19 citations
