Still More Shades of Null: An Evaluation Suite for Responsible Missing Value Imputation [Experiment, Analysis and Benchmark]
Falaah Arif Khan, Denys Herasymuk, Nazar Protsiv, Julia Stoyanovich
摘要
Data missingness is a practical challenge of sustained interest to the scientific community. In this paper, we present Shades-of-Null, an evaluation suite for responsible missing value imputation. Our work is novel in two ways (i) we model realistic and socially-salient missingness scenarios that go beyond Rubin's classic Missing Completely at Random (MCAR), Missing At Random (MAR) and Missing Not At Random (MNAR) settings, to include multi-mechanism missingness (when different missingness patterns co-exist in the data) and missingness shift (when the missingness mechanism changes between training and test) (ii) we evaluate imputers holistically, based on imputation quality and imputation fairness, as well as on the predictive performance, fairness and stability of the models that are trained and tested on the data post-imputation. We use Shades-of-Null to conduct a large-scale empirical study involving 29,736 experimental pipelines, and find that while there is no single best-performing imputation approach for all missingness types, interesting trade-offs arise between predictive performance, fairness and stability, based on the combination of missingness scenario, imputer choice, and the architecture of the predictive model. We make Shades-of-Null publicly available, to enable researchers to rigorously evaluate missing value imputation methods on a wide range of metrics in plausible and socially meaningful scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 被引用 1,847 次
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 被引用 671 次
- not-MIWAE: Deep Generative Modelling with Missing not at Random DataNiels Bruun Ipsen, Pierre-Alexandre Mattei, Jes FrellsenICLR 2021 · 被引用 81 次
- Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain PredictionsBojan Karlas, Peng Li, Renzhi Wu, Nezihe Merve Gürel 等VLDB 2021 · 被引用 69 次
- Identifiable Generative models for Missing Not at Random Data ImputationChao Ma, Cheng ZhangNeurIPS 2021 · 被引用 56 次
相关 Paper
- Uncovering the Propensity Identification Problem in Debiased RecommendationsHonglei Zhang, Shuyi Wang, Haoxuan Li, Chunyuan Zheng 等ICDE 2024 · 被引用 12 次
- Causality-Based Conformal Imputation Correction with Non-Random Missing LabelsChunyuan Zheng, Xiang Li, Hang Pan, Eric Wang 等KDD 2026
- MIRACLE: Causally-Aware Imputation via Learning Missing Data MechanismsTrent Kyono, Yao Zhang, Alexis Bellot, Mihaela van der SchaarNeurIPS 2021 · 被引用 105 次
- Fairness without Imputation: A Decision Tree Approach for Fair Prediction with Missing ValuesHaewon Jeong, Hao Wang, Flávio P. CalmonAAAI 2022 · 被引用 48 次
- RefiDiff: Progressive Refinement Diffusion for Efficient Missing Data ImputationMd. Atik Ahamed, Qiang Ye, Qiang ChengAAAI 2026
