Learning Large DAGs is Harder than you Think: Many Losses are Minimal for the Wrong DAG
Jonas Seng, Matej Zecevic, Devendra Singh Dhami, Kristian Kersting
Abstract
Structure learning is a crucial task in science, especially in fields such as medicine and biology, where the wrong identification of (in)dependencies among random variables can have significant implications. The primary objective of structure learning is to learn a Directed Acyclic Graph (DAG) that represents the underlying probability distribution of the data. Many prominent DAG learners rely on least square losses or log-likelihood losses for optimization. It is well-known from regression models that least square losses are heavily influenced by the scale of the variables. Recently it has been demonstrated that the scale of data also affects performance of structure learning algorithms, though with a strong focus on linear 2-node systems and simulated data. Moving beyond these results, we provide conditions under which square-based losses are minimal for wrong DAGs in ddimensional cases. Furthermore, we also show that scale can impair performance of structure learners if relations among variables are non-linear for both square based and log-likelihood based losses. We confirm our theoretical findings through extensive experiments on synthetic and real-world data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3f7c9dad-212d-4e37-a610-b1341d63d0d4Cited by top-tier papers3
- Markov Equivalence and Consistency in Differentiable Structure LearningChang Deng, Kevin Bello, Pradeep Ravikumar, Bryon AragamNeurIPS 2024 · 8 citations
- Standardizing Structural Causal ModelsWeronika Ormaniec, Scott Sussex, Lars Lorch, Bernhard Schölkopf et al.ICLR 2025
- Differentiable Structure Learning and Causal Discovery for General Binary DataChang Deng, Bryon AragamNeurIPS 2025
Builds on5
- Gradient-Based Neural DAG LearningSébastien Lachapelle, Philippe Brouillard, Tristan Deleu, Simon Lacoste-JulienICLR 2020 · 337 citations
- Beware of the Simulated DAG! Causal Discovery Benchmarks May Be Easy to GameAlexander G. Reisach, Christof Seiler, Sebastian WeichwaldNeurIPS 2021 · 213 citations
- DiBS: Differentiable Bayesian Structure LearningLars Lorch, Jonas Rothfuss, Bernhard Schölkopf, Andreas KrauseNeurIPS 2021 · 144 citations
- DAGs with No Fears: A Closer Look at Continuous Optimization for Learning Bayesian NetworksDennis Wei, Tian Gao, Yue YuNeurIPS 2020 · 102 citations
- Truncated Matrix Power Iteration for Differentiable DAG LearningZhen Zhang, Ignavier Ng, Dong Gong, Yuhang Liu et al.NeurIPS 2022 · 36 citations
Related papers
- On the Role of Sparsity and DAG Constraints for Learning Linear DAGsIgnavier Ng, AmirEmad Ghassami, Kun ZhangNeurIPS 2020 · 306 citations
- Learning Large DAGs by Combining Continuous Optimization and Feedback Arc Set HeuristicsPierre Gillot, Pekka ParviainenAAAI 2022 · 5 citations
- Effective Causal Discovery under Identifiable Heteroscedastic Noise ModelNaiyu Yin, Tian Gao, Yue Yu, Qiang JiAAAI 2024 · 5 citations
- ProDAG: Projected Variational Inference for Directed Acyclic GraphsRyan Thompson, Edwin V. Bonilla, Robert KohnNeurIPS 2025 · 6 citations
- CASTLE: Regularization via Auxiliary Causal Graph DiscoveryTrent Kyono, Yao Zhang, Mihaela van der SchaarNeurIPS 2020 · 82 citations
