ZIP: Lazy Imputation during Query Processing
Yiming Lin, Sharad Mehrotra
Abstract
This paper develops a query-time missing value imputation framework, entitled ZIP, that modifies relational operators to be imputation aware in order to minimize the joint cost of imputing and query processing. The modified operators use a cost-based decision function to determine whether to invoke imputation or to defer to downstream operators to resolve missing values. The modified query processing logic ensures results with deferred imputations are identical to those produced if all missing values were imputed first. ZIP includes a novel outer-join based approach to preserve missing values during execution, and a bloom filter based index to optimize the space and running overhead. Extensive experiments on both real and synthetic data sets demonstrate 10 to 25 times improvement when augmenting the state-of-the-art technology, ImputeDB, with ZIP-based deferred imputation. ZIP also outperforms the offline approach by up to 19607 times in a real data set.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Deduplicated Sampling On-DemandLuca Zecchini, Vasilis Efthymiou, Felix Naumann, Giovanni SimoniniVLDB 2025 · 2 citations
- Hardware-Efficient Data Imputation through DBMS ExtensibilityHubert Mohr-Daurat, Georgios Theodorakis, Holger PirkVLDB 2024 · 2 citations
- DIM-SUM: Dynamic IMputation for Smart Utility ManagementRyan Hildebrant, Rahul Atul Bhope, Sharad Mehrotra, Christopher Tull et al.VLDB 2025 · 2 citations
Builds on8
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang et al.ICDE 2021 · 127 citations
- Missing Value Imputation on Multidimensional Time SeriesParikshit Bansal, Prathamesh Deshpande, Sunita SarawagiVLDB 2021 · 90 citations
- Mind the Gap: An Experimental Evaluation of Imputation of Missing Values Techniques in Time SeriesMourad Khayati, Alberto Lerner, Zakhar Tymchenko, Philippe Cudré-MaurouxVLDB 2020 · 57 citations
- Efficient and Effective Data Imputation with Influence FunctionsXiaoye Miao, Yangyang Wu, Lu Chen, Yunjun Gao et al.VLDB 2022 · 38 citations
- Adaptive Data Augmentation for Supervised Learning over Missing DataTongyu Liu, Ju Fan, Yinqing Luo, Nan Tang et al.VLDB 2021 · 31 citations
Related papers
- In-Database Data ImputationMassimo Perini, Milos NikolicSIGMOD 2024 · 5 citations
- Rethinking Time-Series Imputation as Conditional Inference along Temporal EvolutionYu Fan, Yang Yang, guo yufan, Huazhong Yang et al.ICML 2026
- Pushing Data-Induced Predicates Through Joins in Big-Data ClustersLaurel J. Orr, Srikanth Kandula, Surajit ChaudhuriVLDB 2020 · 35 citations
- Imputing Various Incomplete Attributes via Distance Likelihood MaximizationShaoxu Song, Yu SunKDD 2020 · 15 citations
- Think Twice Before Imputation: Optimizing Data Imputation Order for Machine LearningJiaxuan Zhang, Haitao Yuan, Jianing Si, Nan Jiang et al.ICDE 2025
