Automatic Data Repair: Are We Ready to Deploy?
Wei Ni, Xiaoye Miao, Xiangyu Zhao, Yangyang Wu, Shuwei Liang, Jianwei Yin
Abstract
Data quality is paramount in today's data-driven world, especially in the era of generative AI. Dirty data with errors and inconsistencies usually leads to flawed insights, unreliable decision-making, and biased or low-quality outputs from generative models. The study of repairing erroneous data has gained significant importance. Existing data repair algorithms differ in information utilization, problem settings, and are tested in limited scenarios. In this paper, we compare and summarize these algorithms with a driven information-based taxonomy. We systematically conduct a comprehensive evaluation of 12 mainstream data repair algorithms on 12 datasets under the settings of various data error rates, error types, and 4 downstream analysis tasks, assessing their error reduction performance with a novel but practical metric. We develop an effective and unified repair optimization strategy that substantially benefits the state of the arts. We conclude that, it is always worthy of data repair. The clean data does not determine the upper bound of data analysis performance. We provide valuable guidelines, challenges, and promising directions in the data repair domain. We anticipate this paper enabling researchers and users to well understand and deploy data repair algorithms in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- ZeroED: Hybrid Zero-Shot Error Detection Through Large Language Model ReasoningWei Ni, Kaihang Zhang, Xiaoye Miao, Xiangyu Zhao et al.ICDE 2025 · 5 citations
- Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in TablesQixu Chen, Yeye He, Raymond Chi-Wing Wong, Weiwei Cui et al.SIGMOD 2025 · 4 citations
- UniClean: A Scalable Data Cleaning Solution for Mixed Errors based on Unified Cleaners and Optimized Cleaning WorkflowXiaoou Ding, Zekai Qian, Hongzhi Wang, Siying Chen et al.VLDB 2025 · 1 citation
- ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic WorkflowsWei Liu, Yang Gu, Xi Yan, Zihan Nan et al.KDD 2026 · 1 citation
- Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language ModelsYurong Liu, Yeye He, Haoyu Dong, Junjie Xing et al.VLDB 2026
Builds on14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan et al.VLDB 2021 · 484 citations
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 331 citations
- Generative Semi-supervised Learning for Multivariate Time Series ImputationXiaoye Miao, Yangyang Wu, Jun Wang, Yunjun Gao et al.AAAI 2021 · 212 citations
Related papers
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang et al.ICDE 2021 · 127 citations
- TORepair: Diffusion-Based Task-Oriented Error Repair Via Differentiable Bi-Level OptimizationWei Ni, Xiaoye Miao, Xiangyu Zhao, Yangyang Wu et al.ICDE 2026
- DeMix: Debugging Training Data with Mixed Data Error Types by Investigating Influence VectorsJiale Deng, Yanyan Shen, Xiaogang Shi, Junjun ChaiKDD 2026
- Knowledge-Enhanced Program Repair for Data Science CodeShuyin Ouyang, Jie M. Zhang, Zeyu Sun, Albert Meroño-PeñuelaICSE 2025 · 2 citations
- Generalizable Data Cleaning of Tabular Data in Latent SpaceEduardo Souza dos Reis, Mohamed Abdelaal, Carsten BinnigVLDB 2024 · 3 citations
