DAGs with No Fears: A Closer Look at Continuous Optimization for Learning Bayesian Networks
Dennis Wei, Tian Gao, Yue Yu
Abstract
This paper re-examines a continuous optimization framework dubbed NOTEARS for learning Bayesian networks. We first generalize existing algebraic characterizations of acyclicity to a class of matrix polynomials. Next, focusing on a one-parameter-per-edge setting, it is shown that the Karush-Kuhn-Tucker (KKT) optimality conditions for the NOTEARS formulation cannot be satisfied except in a trivial case, which explains a behavior of the associated algorithm. We then derive the KKT conditions for an equivalent reformulation, show that they are indeed necessary, and relate them to explicit constraints that certain edges be absent from the graph. If the score function is convex, these KKT conditions are also sufficient for local minimality despite the non-convexity of the constraint. Informed by the KKT conditions, a local search post-processing algorithm is proposed and shown to substantially and universally improve the structural Hamming distance of all tested algorithms, typically by a factor of 2 or more. Some combinations with local search are both more accurate and more efficient than the original NOTEARS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ba3af36-6549-406a-a5f6-dcf5c44609a6Cited by top-tier papers38
- DAGMA: Learning DAGs via M-matrices and a Log-Determinant Acyclicity CharacterizationKevin Bello, Bryon Aragam, Pradeep RavikumarNeurIPS 2022 · 222 citations
- Beware of the Simulated DAG! Causal Discovery Benchmarks May Be Easy to GameAlexander G. Reisach, Christof Seiler, Sebastian WeichwaldNeurIPS 2021 · 213 citations
- DiBS: Differentiable Bayesian Structure LearningLars Lorch, Jonas Rothfuss, Bernhard Schölkopf, Andreas KrauseNeurIPS 2021 · 144 citations
- BCD Nets: Scalable Variational Approaches for Bayesian Causal DiscoveryChris Cundy, Aditya Grover, Stefano ErmonNeurIPS 2021 · 105 citations
- Truncated Matrix Power Iteration for Differentiable DAG LearningZhen Zhang, Ignavier Ng, Dong Gong, Yuhang Liu et al.NeurIPS 2022 · 36 citations
Builds on1
Related papers
- Reliable Causal Discovery with Improved Exact Search and Weaker AssumptionsIgnavier Ng, Yujia Zheng, Jiji Zhang, Kun ZhangNeurIPS 2021 · 35 citations
- DAGs with No Curl: An Efficient DAG Structure Learning ApproachYue Yu, Tian Gao, Naiyu Yin, Qiang JiICML 2021 · 77 citations
- Differentiable and Transportable Structure LearningJeroen Berrevoets, Nabeel Seedat, Fergus Imrie, Mihaela van der SchaarICML 2023 · 4 citations
- Learning Large DAGs by Combining Continuous Optimization and Feedback Arc Set HeuristicsPierre Gillot, Pekka ParviainenAAAI 2022 · 5 citations
- DAG Learning on the PermutahedronValentina Zantedeschi, Luca Franceschi, Jean Kaddour, Matt J. Kusner et al.ICLR 2023
