Efficient and Effective Data Imputation with Influence Functions
Xiaoye Miao, Yangyang Wu, Lu Chen, Yunjun Gao, Jun Wang, Jianwei Yin
摘要
Data imputation has been extensively explored to solve the missing data problem. The dramatically rising volume of missing data makes the training of imputation models computationally infeasible in real-life scenarios. In this paper, we propose an efficient and effective data imputation system with influence functions , named EDIT, which quickly trains a parametric imputation model with representative samples under imputation accuracy guarantees. EDIT mainly consists of two modules, i.e., an imputation influence evaluation (IIE) module and a representative sample selection (RSS) module. IIE leverages the influence functions to estimate the effect of (in)complete samples on the prediction result of parametric imputation models. RSS builds a minimum set of the high-effect samples to satisfy a user-specified imputation accuracy. Moreover, we introduce a weighted loss function that drives the parametric imputation model to pay more attention on the high-effect samples. Extensive experiments upon ten state-of-the-art imputation methods demonstrate that, EDIT adopts only about 5% samples to speed up the model training by 4x in average with more than 11% accuracy gain.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Frequency-aware Generative Models for Multivariate Time Series ImputationXinyu Yang, Yu Sun, Xiaojie Yuan, Xinyang ChenNeurIPS 2024 · 被引用 41 次
- GoodCore: Data-effective and Data-efficient Machine Learning through Coreset Selection over Incomplete DataChengliang Chai, Jiabin Liu, Nan Tang, Ju Fan 等SIGMOD 2023 · 被引用 37 次
- Automatic Data Repair: Are We Ready to Deploy?Wei Ni, Xiaoye Miao, Xiangyu Zhao, Yangyang Wu 等VLDB 2024 · 被引用 26 次
- Outlier Summarization via Human Interpretable RulesYuhao Deng, Yu Wang, Lei Cao, Lianpeng Qiao 等VLDB 2024 · 被引用 6 次
- In-Database Data ImputationMassimo Perini, Milos NikolicSIGMOD 2024 · 被引用 5 次
它引用的顶会 Paper5
- Generative Semi-supervised Learning for Multivariate Time Series ImputationXiaoye Miao, Yangyang Wu, Jun Wang, Yunjun Gao 等AAAI 2021 · 被引用 212 次
- Missing Data Imputation using Optimal TransportBoris Muzellec, Julie Josse, Claire Boyer, Marco CuturiICML 2020 · 被引用 179 次
- Not All Unlabeled Data are Equal: Learning to Weight Data in Semi-supervised LearningZhongzheng Ren, Raymond A. Yeh, Alexander G. SchwingNeurIPS 2020 · 被引用 106 次
- Adaptive Data Augmentation for Supervised Learning over Missing DataTongyu Liu, Ju Fan, Yinqing Luo, Nan Tang 等VLDB 2021 · 被引用 31 次
- Baran: Effective Error Correction via a Unified Context Representation and Transfer LearningMohammad Mahdavi, Ziawasch AbedjanVLDB 2020
相关 Paper
- Towards a statistical theory of data selection under weak supervisionGermain Kolossov, Andrea Montanari, Pulkit TandonICLR 2024 · 被引用 27 次
- Certain and Approximately Certain Models for Statistical LearningCheng Zhen, Nischal Aryal, Arash Termehchy, Amandeep Singh ChabadaSIGMOD 2024 · 被引用 4 次
- Resolving Training Biases via Influence-based Data RelabelingShuming Kong, Yanyan Shen, Linpeng HuangICLR 2022 · 被引用 71 次
- FastIF: Scalable Influence Functions for Efficient Model Interpretation and DebuggingHan Guo, Nazneen Rajani, Peter Hase, Mohit Bansal 等EMNLP 2021 · 被引用 51 次
- Data Glitches Discovery using Influence-based Model ExplanationsNikolaos Myrtakis, Ioannis Tsamardinos, Vassilis ChristophidesKDD 2025
