In-Database Data Imputation
Massimo Perini, Milos Nikolic
Abstract
Missing data is a widespread problem in many domains, creating challenges in data analysis and decision making. Traditional techniques for dealing with missing data, such as excluding incomplete records or imputing simple estimates (e.g., mean), are computationally efficient but may introduce bias and disrupt variable relationships, leading to inaccurate analyses. Model-based imputation techniques offer a more robust solution that preserves the variability and relationships in the data, but they demand significantly more computation time, limiting their applicability to small datasets.
This work enables efficient, high-quality, and scalable data imputation within a database system using the widely used MICE method. We adapt this method to exploit computation sharing and a ring abstraction for faster model training. To impute both continuous and categorical values, we develop techniques for in-database learning of stochastic linear regression and Gaussian discriminant analysis models. Our MICE implementations in PostgreSQL and DuckDB outperform alternative MICE implementations and model-based imputation techniques by up to two orders of magnitude in terms of computation time, while maintaining high imputation quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 44a3efea-d625-43f4-bfe3-0775dfab2be2Cited by top-tier papers1
Ask how each one uses itBuilds on19
- Handling Missing Data with Graph Representation LearningJiaxuan You, Xiaobai Ma, Daisy Yi Ding, Mykel J. Kochenderfer et al.NeurIPS 2020 · 274 citations
- HyperImpute: Generalized Iterative Imputation with Automatic Model SelectionDaniel Jarrett, Bogdan Cebere, Tennison Liu, Alicia Curth et al.ICML 2022 · 129 citations
- MIRACLE: Causally-Aware Imputation via Learning Missing Data MechanismsTrent Kyono, Yao Zhang, Alexis Bellot, Mihaela van der SchaarNeurIPS 2021 · 105 citations
- Horizon: Scalable Dependency-driven Data CleaningEl Kindi Rezig, Mourad Ouzzani, Walid G. Aref, Ahmed K. Elmagarmid et al.VLDB 2021 · 95 citations
- Missing Value Imputation on Multidimensional Time SeriesParikshit Bansal, Prathamesh Deshpande, Sunita SarawagiVLDB 2021 · 90 citations
Related papers
- RefiDiff: Progressive Refinement Diffusion for Efficient Missing Data ImputationMd. Atik Ahamed, Qiang Ye, Qiang ChengAAAI 2026
- Iterative Missing Data Imputation with Model Form Adaptation and Non-Missing Feature SupervisionHao Wang, Zhengnan Li, Zhichao Chen, Xu Chen et al.NeurIPS 2025 · 12 citations
- Missing Value Imputation for Mixed Data via Gaussian CopulaYuxuan Zhao, Madeleine UdellKDD 2020 · 53 citations
- ZIP: Lazy Imputation during Query ProcessingYiming Lin, Sharad MehrotraVLDB 2024 · 4 citations
- Certain and Approximately Certain Models for Statistical LearningCheng Zhen, Nischal Aryal, Arash Termehchy, Amandeep Singh ChabadaSIGMOD 2024 · 4 citations
