MissScore: High-Order Score Estimation in the Presence of Missing Data
Wenqin Liu, Haoze Hou, Erdun Gao, Biwei Huang, Qiuhong Ke, Howard D. Bondell, Mingming Gong
摘要
The first order derivative (score) of data density, typically estimated via denoising score matching, has emerged as an effective tool for modeling data distribution and generating synthetic data. Extending this concept to higher-order scores could uncover more detailed local information of the data distribution, enabling new applications. However, learning these high-order scores usually requires complete data, which is often unavailable in real-world scenarios such as healthcare and finance due to privacy and cost constraints. In this work, we introduce MissScore, a novel score-based framework for learning high-order scores from observations with missing data. We derive objective functions for estimating high-order scores under different missing data mechanisms and propose a new algorithm to handle missing data effectively. Our empirical results demonstrate that MissScore efficiently and accurately approximates high-order scores with missing data, while enhancing sampling speed and data quality, as validated through several downstream tasks, including data generation and causal discovery. Under review as a conference paper at ICLR 2025 Theorem 3 If the missing mechanism of x is MCAR, with the missing probability of every element lying between 0 and 1, i.e., p(m i = 1) ∈ [0, 1) for all i ∈ 1, 2 . . . , d. We denote the objective J DSM (θ) = E x,m E x|x,m s 1 (x; θ) + 1 σ 2 (x -x) ⊙ (1 -m)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Improved Techniques for Training Score-Based Generative ModelsYang Song, Stefano ErmonNeurIPS 2020 · 被引用 1,527 次
- Solving Inverse Problems in Medical Imaging with Score-Based Generative ModelsYang Song, Liyue Shen, Lei Xing, Stefano ErmonICLR 2022 · 被引用 721 次
- DAGMA: Learning DAGs via M-matrices and a Log-Determinant Acyclicity CharacterizationKevin Bello, Bryon Aragam, Pradeep RavikumarNeurIPS 2022 · 被引用 222 次
- Missing Data Imputation using Optimal TransportBoris Muzellec, Julie Josse, Claire Boyer, Marco CuturiICML 2020 · 被引用 179 次
相关 Paper
- Estimating High Order Gradients of the Data Distribution by DenoisingChenlin Meng, Yang Song, Wenzhe Li, Stefano ErmonNeurIPS 2021 · 被引用 85 次
- Maximum Likelihood Training for Score-based Diffusion ODEs by High Order Denoising Score MatchingCheng Lu, Kaiwen Zheng, Fan Bao, Jianfei Chen 等ICML 2022 · 被引用 109 次
- Optimal Transport for Structure Learning Under Missing DataVy Vo, He Zhao, Trung Le, Edwin V. Bonilla 等ICML 2024 · 被引用 6 次
- Score Matching with Missing DataJosh Givens, Song Liu, Henry W. J. ReeveICML 2025
- Efficient Learning of Generative Models via Finite-Difference Score MatchingTianyu Pang, Taufik Xu, Chongxuan Li, Yang Song 等NeurIPS 2020 · 被引用 67 次
