Why Not to Use Zero Imputation? Correcting Sparsity Bias in Training Neural Networks
Joonyoung Yi, Juhyuk Lee, Kwang Joon Kim, Sung Ju Hwang, Eunho Yang
Abstract
Handling missing data is one of the most fundamental problems in machine learning. Among many approaches, the simplest and most intuitive way is zero imputation, which treats the value of a missing entry simply as zero. However, many studies have experimentally confirmed that zero imputation results in suboptimal performances in training neural networks. Yet, none of the existing work has explained what brings such performance degradations. In this paper, we introduce the variable sparsity problem (VSP), which describes a phenomenon where the output of a predictive model largely varies with respect to the rate of missingness in the given input, and show that it adversarially affects the model performance. We first theoretically analyze this phenomenon and propose a simple yet effective technique to handle missingness, which we refer to as Sparsity Normalization (SN), that directly targets and resolves the VSP. We further experimentally validate SN on diverse benchmark datasets, to show that debiasing the effect of input-level sparsity improves the performance and stabilizes the training of neural networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f26da744-fd43-4e27-af83-116b9204f2c6Cited by top-tier papers5
- How to deal with missing data in supervised deep learning?Niels Bruun Ipsen, Pierre-Alexandre Mattei, Jes FrellsenICLR 2022 · 39 citations
- Debiasing Averaged Stochastic Gradient Descent to handle missing valuesAude Sportisse, Claire Boyer, Aymeric Dieuleveut, Julie JosseNeurIPS 2020 · 12 citations
- Neuron-Enhanced AutoEncoder Matrix Completion and Collaborative Filtering: Theory and PracticeJicong Fan, Rui Chen, Zhao Zhang, Chris DingICLR 2024 · 4 citations
- Uncovering the Missing Pattern: Unified Framework Towards Trajectory Imputation and PredictionYi Xu, Armin Bazarjani, Hyung-Gun Chi, Chiho Choi et al.CVPR 2023
- AugMask: Training Diffusion Models on Incomplete Tabular Data via Stochastic Augmentation and MaskingJungkyu Kim, Taeyoung Park, Kibok LeeICML 2026
Related papers
- Naive imputation implicitly regularizes high-dimensional linear modelsAlexis Ayme, Claire Boyer, Aymeric Dieuleveut, Erwan ScornetICML 2023 · 10 citations
- Handling Missing Data with Graph Representation LearningJiaxuan You, Xiaobai Ma, Daisy Yi Ding, Mykel J. Kochenderfer et al.NeurIPS 2020 · 274 citations
- What's a good imputation to predict with missing values?Marine Le Morvan, Julie Josse, Erwan Scornet, Gaël VaroquauxNeurIPS 2021 · 95 citations
- Imputation for prediction: beware of diminishing returnsMarine Le Morvan, Gaël VaroquauxICLR 2025
- Random features models: a way to study the success of naive imputationAlexis Ayme, Claire Boyer, Aymeric Dieuleveut, Erwan ScornetICML 2024 · 7 citations
