Automatic Data Repair: Are We Ready to Deploy?
Wei Ni, Xiaoye Miao, Xiangyu Zhao, Yangyang Wu, Shuwei Liang, Jianwei Yin
摘要
Data quality is paramount in today's data-driven world, especially in the era of generative AI. Dirty data with errors and inconsistencies usually leads to flawed insights, unreliable decision-making, and biased or low-quality outputs from generative models. The study of repairing erroneous data has gained significant importance. Existing data repair algorithms differ in information utilization, problem settings, and are tested in limited scenarios. In this paper, we compare and summarize these algorithms with a driven information-based taxonomy. We systematically conduct a comprehensive evaluation of 12 mainstream data repair algorithms on 12 datasets under the settings of various data error rates, error types, and 4 downstream analysis tasks, assessing their error reduction performance with a novel but practical metric. We develop an effective and unified repair optimization strategy that substantially benefits the state of the arts. We conclude that, it is always worthy of data repair. The clean data does not determine the upper bound of data analysis performance. We provide valuable guidelines, challenges, and promising directions in the data repair domain. We anticipate this paper enabling researchers and users to well understand and deploy data repair algorithms in practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- ZeroED: Hybrid Zero-Shot Error Detection Through Large Language Model ReasoningWei Ni, Kaihang Zhang, Xiaoye Miao, Xiangyu Zhao 等ICDE 2025 · 被引用 5 次
- Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in TablesQixu Chen, Yeye He, Raymond Chi-Wing Wong, Weiwei Cui 等SIGMOD 2025 · 被引用 4 次
- UniClean: A Scalable Data Cleaning Solution for Mixed Errors based on Unified Cleaners and Optimized Cleaning WorkflowXiaoou Ding, Zekai Qian, Hongzhi Wang, Siying Chen 等VLDB 2025 · 被引用 1 次
- ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic WorkflowsWei Liu, Yang Gu, Xi Yan, Zihan Nan 等KDD 2026 · 被引用 1 次
- Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language ModelsYurong Liu, Yeye He, Haoyu Dong, Junjie Xing 等VLDB 2026
它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan 等VLDB 2021 · 被引用 484 次
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 被引用 331 次
- Generative Semi-supervised Learning for Multivariate Time Series ImputationXiaoye Miao, Yangyang Wu, Jun Wang, Yunjun Gao 等AAAI 2021 · 被引用 212 次
相关 Paper
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang 等ICDE 2021 · 被引用 127 次
- TORepair: Diffusion-Based Task-Oriented Error Repair Via Differentiable Bi-Level OptimizationWei Ni, Xiaoye Miao, Xiangyu Zhao, Yangyang Wu 等ICDE 2026
- DeMix: Debugging Training Data with Mixed Data Error Types by Investigating Influence VectorsJiale Deng, Yanyan Shen, Xiaogang Shi, Junjun ChaiKDD 2026
- Knowledge-Enhanced Program Repair for Data Science CodeShuyin Ouyang, Jie M. Zhang, Zeyu Sun, Albert Meroño-PeñuelaICSE 2025 · 被引用 2 次
- Generalizable Data Cleaning of Tabular Data in Latent SpaceEduardo Souza dos Reis, Mohamed Abdelaal, Carsten BinnigVLDB 2024 · 被引用 3 次
