Optimal Transport for Structure Learning Under Missing Data
Vy Vo, He Zhao, Trung Le, Edwin V. Bonilla, Dinh Phung
摘要
Causal discovery in the presence of missing data introduces a chicken-and-egg dilemma. While the goal is to recover the true causal structure, robust imputation requires considering the dependencies or, preferably, causal relations among variables. Merely filling in missing values with existing imputation methods and subsequently applying structure learning on the complete data is empirically shown to be sub-optimal. To address this problem, we propose a score-based algorithm for learning causal structures from missing data based on optimal transport. This optimal transport viewpoint diverges from existing score-based approaches that are dominantly based on expectation maximization. We formulate structure learning as a density fitting problem, where the goal is to find the causal model that induces a distribution of minimum Wasserstein distance with the observed data distribution. Our framework is shown to recover the true causal graphs more effectively than competing methods in most simulations and real-data settings. Empirical evidence also shows the superior scalability of our approach, along with the flexibility to incorporate any off-the-shelf causal discovery methods for complete data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Distribution Alignment Optimization through Neural Collapse for Long-tailed ClassificationJintong Gao, He Zhao, Dandan Guo, Hongyuan ZhaICML 2024 · 被引用 27 次
- Neural Topic Modeling with Large Language Models in the LoopXiaohao Yang, He Zhao, Weijie Xu, Yuanyuan Qi 等ACL 2025 · 被引用 13 次
- Parameter Estimation in DAGs from Incomplete Data via Optimal TransportVy Vo, Trung Le, Long Tung Vuong, He Zhao 等ICML 2024 · 被引用 5 次
- Ordering-based Causal Discovery via Generalized Score MatchingVy Vo, Trung Le, He Zhao, Edwin V. Bonilla 等KDD 2026 · 被引用 1 次
- Robust Simulation-Based Inference under Missing Data via Neural ProcessesYogesh Verma, Ayush Bharti, Vikas GargICLR 2025
它引用的顶会 Paper27
- Gradient-Based Neural DAG LearningSébastien Lachapelle, Philippe Brouillard, Tristan Deleu, Simon Lacoste-JulienICLR 2020 · 被引用 337 次
- On the Role of Sparsity and DAG Constraints for Learning Linear DAGsIgnavier Ng, AmirEmad Ghassami, Kun ZhangNeurIPS 2020 · 被引用 306 次
- Causal Discovery with Reinforcement LearningShengyu Zhu, Ignavier Ng, Zhitang ChenICLR 2020 · 被引用 285 次
- Handling Missing Data with Graph Representation LearningJiaxuan You, Xiaobai Ma, Daisy Yi Ding, Mykel J. Kochenderfer 等NeurIPS 2020 · 被引用 274 次
- Missing Data Imputation using Optimal TransportBoris Muzellec, Julie Josse, Claire Boyer, Marco CuturiICML 2020 · 被引用 179 次
相关 Paper
- MissDAG: Causal Discovery in the Presence of Missing Data with Continuous Additive Noise ModelsErdun Gao, Ignavier Ng, Mingming Gong, Li Shen 等NeurIPS 2022 · 被引用 36 次
- MissScore: High-Order Score Estimation in the Presence of Missing DataWenqin Liu, Haoze Hou, Erdun Gao, Biwei Huang 等ICML 2025
- Score Matching Enables Causal Discovery of Nonlinear Additive Noise ModelsPaul Rolland, Volkan Cevher, Matthäus Kleindessner, Chris Russell 等ICML 2022 · 被引用 123 次
- Learning to Induce Causal StructureNan Rosemary Ke, Silvia Chiappa, Jane X. Wang, Jörg Bornschein 等ICLR 2023 · 被引用 17 次
- Causal Discovery for Irregularly Time Series with Consistency GuaranteesWeihong Li, Baohong Li, Anpeng Wu, Zhihan Li 等ICML 2026
