MissScore: High-Order Score Estimation in the Presence of Missing Data
Wenqin Liu, Haoze Hou, Erdun Gao, Biwei Huang, Qiuhong Ke, Howard D. Bondell, Mingming Gong
Abstract
The first order derivative (score) of data density, typically estimated via denoising score matching, has emerged as an effective tool for modeling data distribution and generating synthetic data. Extending this concept to higher-order scores could uncover more detailed local information of the data distribution, enabling new applications. However, learning these high-order scores usually requires complete data, which is often unavailable in real-world scenarios such as healthcare and finance due to privacy and cost constraints. In this work, we introduce MissScore, a novel score-based framework for learning high-order scores from observations with missing data. We derive objective functions for estimating high-order scores under different missing data mechanisms and propose a new algorithm to handle missing data effectively. Our empirical results demonstrate that MissScore efficiently and accurately approximates high-order scores with missing data, while enhancing sampling speed and data quality, as validated through several downstream tasks, including data generation and causal discovery. Under review as a conference paper at ICLR 2025 Theorem 3 If the missing mechanism of x is MCAR, with the missing probability of every element lying between 0 and 1, i.e., p(m i = 1) ∈ [0, 1) for all i ∈ 1, 2 . . . , d. We denote the objective J DSM (θ) = E x,m E x|x,m s 1 (x; θ) + 1 σ 2 (x -x) ⊙ (1 -m)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a71e32af-f5f7-4e4a-98b5-0de0cbaf6357Builds on12
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Improved Techniques for Training Score-Based Generative ModelsYang Song, Stefano ErmonNeurIPS 2020 · 1,527 citations
- Solving Inverse Problems in Medical Imaging with Score-Based Generative ModelsYang Song, Liyue Shen, Lei Xing, Stefano ErmonICLR 2022 · 721 citations
- DAGMA: Learning DAGs via M-matrices and a Log-Determinant Acyclicity CharacterizationKevin Bello, Bryon Aragam, Pradeep RavikumarNeurIPS 2022 · 222 citations
- Missing Data Imputation using Optimal TransportBoris Muzellec, Julie Josse, Claire Boyer, Marco CuturiICML 2020 · 179 citations
Related papers
- Estimating High Order Gradients of the Data Distribution by DenoisingChenlin Meng, Yang Song, Wenzhe Li, Stefano ErmonNeurIPS 2021 · 85 citations
- Maximum Likelihood Training for Score-based Diffusion ODEs by High Order Denoising Score MatchingCheng Lu, Kaiwen Zheng, Fan Bao, Jianfei Chen et al.ICML 2022 · 109 citations
- Optimal Transport for Structure Learning Under Missing DataVy Vo, He Zhao, Trung Le, Edwin V. Bonilla et al.ICML 2024 · 6 citations
- Score Matching with Missing DataJosh Givens, Song Liu, Henry W. J. ReeveICML 2025
- Efficient Learning of Generative Models via Finite-Difference Score MatchingTianyu Pang, Taufik Xu, Chongxuan Li, Yang Song et al.NeurIPS 2020 · 67 citations
