Improving the imputation of missing data with Markov Blanket discovery
Yang Liu, Anthony C. Constantinou
摘要
The process of imputation of missing data typically relies on generative and regression models. These approaches often operate on the unrealistic assumption that all of the data features are directly related with one another, and use all of the available features to impute missing values. In this paper, we propose a novel Markov Blanket discovery approach to determine the optimal feature set for a given variable by considering both observed variables and missingness of partially observed variables to account for systematic missingness. We then incorporate this method to the learning process of the state-of-the-art MissForest imputation algorithm, such that it informs MissForest which features to consider to impute missing values, depending on the variable the missing value belongs to. Experiments across different case studies and multiple imputation algorithms show that the proposed solution improves imputation accuracy, both under random and systematic missingness. Recently, causal information has also been adopted to feature selection for missing data imputation. Kyono et al. (2021) proposed to impute missing values of a variable given its causal parents derived from the weights of the input layer in the neural network. Similarly, Yu et al. (2022) proposed the MimMB framework that learns Markov Blankets (MBs) to be used for feature selection in imputation, which is an iterative process that learns MBs from the imputed data and updates the learned MB after each iteration. Note that while MimMB is related to our work, since we also use MB construction for feature selection, an important distinction between the two is that MimMB combines MBs with imputed data whereas, as we later describe in Section 3, the learning phase of MBs that we propose is separated from imputation, accounts for partially observed variables, and improves computational efficiency. In this paper, we use the graphical expression of missingness proposed by Mohan et al. (2013) , known as m-graph, which is a graph that captures observed variables in conjunction with the possible
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- Handling Missing Data with Graph Representation LearningJiaxuan You, Xiaobai Ma, Daisy Yi Ding, Mykel J. Kochenderfer 等NeurIPS 2020 · 被引用 274 次
- Missing Data Imputation using Optimal TransportBoris Muzellec, Julie Josse, Claire Boyer, Marco CuturiICML 2020 · 被引用 179 次
- HyperImpute: Generalized Iterative Imputation with Automatic Model SelectionDaniel Jarrett, Bogdan Cebere, Tennison Liu, Alicia Curth 等ICML 2022 · 被引用 129 次
- MIRACLE: Causally-Aware Imputation via Learning Missing Data MechanismsTrent Kyono, Yao Zhang, Alexis Bellot, Mihaela van der SchaarNeurIPS 2021 · 被引用 105 次
相关 Paper
- Multi-Label Causal Feature SelectionXingyu Wu, Bingbing Jiang, Kui Yu, Huanhuan Chen 等AAAI 2020 · 被引用 68 次
- MissDAG: Causal Discovery in the Presence of Missing Data with Continuous Additive Noise ModelsErdun Gao, Ignavier Ng, Mingming Gong, Li Shen 等NeurIPS 2022 · 被引用 36 次
- Mining of Switching Sparse Networks for Missing Value Imputation in Multivariate Time SeriesKohei Obata, Koki Kawabata, Yasuko Matsubara, Yasushi SakuraiKDD 2024 · 被引用 6 次
- RefiDiff: Progressive Refinement Diffusion for Efficient Missing Data ImputationMd. Atik Ahamed, Qiang Ye, Qiang ChengAAAI 2026
- CACTI: Leveraging Copy Masking and Contextual Information to Improve Tabular Data ImputationAditya Gorla, Ryan Wang, Zhengtong Liu, Ulzee An 等ICML 2025
