Markov Equivalence and Consistency in Differentiable Structure Learning
Chang Deng, Kevin Bello, Pradeep Ravikumar, Bryon Aragam
Abstract
Existing approaches to differentiable structure learning of directed acyclic graphs (DAGs) rely on strong identifiability assumptions in order to guarantee that global minimizers of the acyclicity-constrained optimization problem identifies the true DAG. Moreover, it has been observed empirically that the optimizer may exploit undesirable artifacts in the loss function. We explain and remedy these issues by studying the behavior of differentiable acyclicity-constrained programs under general likelihoods with multiple global minimizers. By carefully regularizing the likelihood, it is possible to identify the sparsest model in the Markov equivalence class, even in the absence of an identifiable parametrization. We first study the Gaussian case in detail, showing how proper regularization of the likelihood defines a score that identifies the sparsest model. Assuming faithfulness, it also recovers the Markov equivalence class. These results are then generalized to general models and likelihoods, where the same claims hold. These theoretical results are validated empirically, showing how this can be done using standard gradient-based optimizers, thus paving the way for differentiable structure learning under general models and losses.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Differentiable Structure Learning with Ancestral ConstraintsTaiyu Ban, Changxin Rong, Xiangyu Wang, Lyuzhou Chen et al.ICML 2025
- Differentiable Structure Learning and Causal Discovery for General Binary DataChang Deng, Bryon AragamNeurIPS 2025
Builds on11
- On the Role of Sparsity and DAG Constraints for Learning Linear DAGsIgnavier Ng, AmirEmad Ghassami, Kun ZhangNeurIPS 2020 · 306 citations
- Differentiable Causal Discovery from Interventional DataPhilippe Brouillard, Sébastien Lachapelle, Alexandre Lacoste, Simon Lacoste-Julien et al.NeurIPS 2020 · 295 citations
- DAGMA: Learning DAGs via M-matrices and a Log-Determinant Acyclicity CharacterizationKevin Bello, Bryon Aragam, Pradeep RavikumarNeurIPS 2022 · 222 citations
- Beware of the Simulated DAG! Causal Discovery Benchmarks May Be Easy to GameAlexander G. Reisach, Christof Seiler, Sebastian WeichwaldNeurIPS 2021 · 213 citations
- DAGs with No Fears: A Closer Look at Continuous Optimization for Learning Bayesian NetworksDennis Wei, Tian Gao, Yue YuNeurIPS 2020 · 102 citations
Related papers
- Revisiting Differentiable Structure Learning: Inconsistency of L1 Penalty and BeyondKaifeng Jin, Ignavier Ng, Kun Zhang, Biwei HuangAAAI 2026
- Constraint-Free Structure Learning with Smooth Acyclic OrientationsRiccardo Massidda, Francesco Landolfi, Martina Cinquini, Davide BacciuICLR 2024 · 10 citations
- CoLiDE: Concomitant Linear DAG EstimationSeyed Saman Saboksayr, Gonzalo Mateos, Mariano TepperICLR 2024 · 9 citations
- Characterizing Distribution Equivalence and Structure Learning for Cyclic and Acyclic Directed GraphsAmirEmad Ghassami, Alan Yang, Negar Kiyavash, Kun ZhangICML 2020 · 32 citations
- DAGs with No Curl: An Efficient DAG Structure Learning ApproachYue Yu, Tian Gao, Naiyu Yin, Qiang JiICML 2021 · 77 citations
