ZIP: Lazy Imputation during Query Processing
Yiming Lin, Sharad Mehrotra
摘要
This paper develops a query-time missing value imputation framework, entitled ZIP, that modifies relational operators to be imputation aware in order to minimize the joint cost of imputing and query processing. The modified operators use a cost-based decision function to determine whether to invoke imputation or to defer to downstream operators to resolve missing values. The modified query processing logic ensures results with deferred imputations are identical to those produced if all missing values were imputed first. ZIP includes a novel outer-join based approach to preserve missing values during execution, and a bloom filter based index to optimize the space and running overhead. Extensive experiments on both real and synthetic data sets demonstrate 10 to 25 times improvement when augmenting the state-of-the-art technology, ImputeDB, with ZIP-based deferred imputation. ZIP also outperforms the offline approach by up to 19607 times in a real data set.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Deduplicated Sampling On-DemandLuca Zecchini, Vasilis Efthymiou, Felix Naumann, Giovanni SimoniniVLDB 2025 · 被引用 2 次
- Hardware-Efficient Data Imputation through DBMS ExtensibilityHubert Mohr-Daurat, Georgios Theodorakis, Holger PirkVLDB 2024 · 被引用 2 次
- DIM-SUM: Dynamic IMputation for Smart Utility ManagementRyan Hildebrant, Rahul Atul Bhope, Sharad Mehrotra, Christopher Tull 等VLDB 2025 · 被引用 2 次
它引用的顶会 Paper8
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang 等ICDE 2021 · 被引用 127 次
- Missing Value Imputation on Multidimensional Time SeriesParikshit Bansal, Prathamesh Deshpande, Sunita SarawagiVLDB 2021 · 被引用 90 次
- Mind the Gap: An Experimental Evaluation of Imputation of Missing Values Techniques in Time SeriesMourad Khayati, Alberto Lerner, Zakhar Tymchenko, Philippe Cudré-MaurouxVLDB 2020 · 被引用 57 次
- Efficient and Effective Data Imputation with Influence FunctionsXiaoye Miao, Yangyang Wu, Lu Chen, Yunjun Gao 等VLDB 2022 · 被引用 38 次
- Adaptive Data Augmentation for Supervised Learning over Missing DataTongyu Liu, Ju Fan, Yinqing Luo, Nan Tang 等VLDB 2021 · 被引用 31 次
相关 Paper
- In-Database Data ImputationMassimo Perini, Milos NikolicSIGMOD 2024 · 被引用 5 次
- Rethinking Time-Series Imputation as Conditional Inference along Temporal EvolutionYu Fan, Yang Yang, guo yufan, Huazhong Yang 等ICML 2026
- Pushing Data-Induced Predicates Through Joins in Big-Data ClustersLaurel J. Orr, Srikanth Kandula, Surajit ChaudhuriVLDB 2020 · 被引用 35 次
- Imputing Various Incomplete Attributes via Distance Likelihood MaximizationShaoxu Song, Yu SunKDD 2020 · 被引用 15 次
- Think Twice Before Imputation: Optimizing Data Imputation Order for Machine LearningJiaxuan Zhang, Haitao Yuan, Jianing Si, Nan Jiang 等ICDE 2025
