Efficient and Effective Data Imputation with Influence Functions
Xiaoye Miao, Yangyang Wu, Lu Chen, Yunjun Gao, Jun Wang, Jianwei Yin
Abstract
Data imputation has been extensively explored to solve the missing data problem. The dramatically rising volume of missing data makes the training of imputation models computationally infeasible in real-life scenarios. In this paper, we propose an efficient and effective data imputation system with influence functions , named EDIT, which quickly trains a parametric imputation model with representative samples under imputation accuracy guarantees. EDIT mainly consists of two modules, i.e., an imputation influence evaluation (IIE) module and a representative sample selection (RSS) module. IIE leverages the influence functions to estimate the effect of (in)complete samples on the prediction result of parametric imputation models. RSS builds a minimum set of the high-effect samples to satisfy a user-specified imputation accuracy. Moreover, we introduce a weighted loss function that drives the parametric imputation model to pay more attention on the high-effect samples. Extensive experiments upon ten state-of-the-art imputation methods demonstrate that, EDIT adopts only about 5% samples to speed up the model training by 4x in average with more than 11% accuracy gain.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d701447f-2170-4082-af71-5bc009c4debfCited by top-tier papers11
- Frequency-aware Generative Models for Multivariate Time Series ImputationXinyu Yang, Yu Sun, Xiaojie Yuan, Xinyang ChenNeurIPS 2024 · 41 citations
- GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete DataChengliang Chai, Jiabin Liu, Nan Tang, Ju Fan et al.SIGMOD 2023 · 37 citations
- Automatic Data Repair: Are We Ready to Deploy?Wei Ni, Xiaoye Miao, Xiangyu Zhao, Yangyang Wu et al.VLDB 2024 · 26 citations
- Outlier Summarization via Human Interpretable RulesYuhao Deng, Yu Wang, Lei Cao, Lianpeng Qiao et al.VLDB 2024 · 6 citations
- In-Database Data ImputationMassimo Perini, Milos NikolicSIGMOD 2024 · 5 citations
Builds on5
- Generative Semi-supervised Learning for Multivariate Time Series ImputationXiaoye Miao, Yangyang Wu, Jun Wang, Yunjun Gao et al.AAAI 2021 · 212 citations
- Missing Data Imputation using Optimal TransportBoris Muzellec, Julie Josse, Claire Boyer, Marco CuturiICML 2020 · 179 citations
- Not All Unlabeled Data are Equal: Learning to Weight Data in Semi-supervised LearningZhongzheng Ren, Raymond A. Yeh, Alexander G. SchwingNeurIPS 2020 · 106 citations
- Adaptive Data Augmentation for Supervised Learning over Missing DataTongyu Liu, Ju Fan, Yinqing Luo, Nan Tang et al.VLDB 2021 · 31 citations
- Baran: Effective Error Correction via a Unified Context Representation and Transfer LearningMohammad Mahdavi, Ziawasch AbedjanVLDB 2020
Related papers
- Towards a statistical theory of data selection under weak supervisionGermain Kolossov, Andrea Montanari, Pulkit TandonICLR 2024 · 27 citations
- Certain and Approximately Certain Models for Statistical LearningCheng Zhen, Nischal Aryal, Arash Termehchy, Amandeep Singh ChabadaSIGMOD 2024 · 4 citations
- Resolving Training Biases via Influence-based Data RelabelingShuming Kong, Yanyan Shen, Linpeng HuangICLR 2022 · 71 citations
- FastIF: Scalable Influence Functions for Efficient Model Interpretation and DebuggingHan Guo, Nazneen Rajani, Peter Hase, Mohit Bansal et al.EMNLP 2021 · 51 citations
- Data Glitches Discovery using Influence-based Model ExplanationsNikolaos Myrtakis, Ioannis Tsamardinos, Vassilis ChristophidesKDD 2025
