UniClean: A Scalable Data Cleaning Solution for Mixed Errors based on Unified Cleaners and Optimized Cleaning Workflow
Xiaoou Ding, Zekai Qian, Hongzhi Wang, Siying Chen, Yafeng Tang, Hongbin Su, Huan Hu, Chen Wang
Abstract
Data cleaning is an essential technique to enhance data quality. Despite the proposal of various algorithms with different cleaning strategies, current automated cleaning technologies still fall short of practical requirements when dealing with large-scale data containing mixed errors. This paper presents UniClean to efficiently solve the mixed error cleaning problem with three key technical contributions. (1) A unified construction and extension method for cleaners, enabling cleaning methods to easily utilize various cleaners to perform cleaning tasks. (2) Three optimization strategies to achieve efficiency-oriented cleaning preparation. (3) A cleaning algorithm based on an optimized cleaning process to effectively clean mixed errors. UniClean achieves a time complexity of O (| D error | 4 · | Op | + |D| · | D error |), significantly enhancing scalability. Experiments on public and large-scale enterprise datasets demonstrate that UniClean achieves over 40% improvement across five metrics, compared to five state-of-the-art cleaning methods, and delivers more than 30% gains in F1 and REDR on complex datasets, while completing the cleaning process within hours even for millions of records.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d771720-5f73-4bb5-ad69-a91e59689971Cited by top-tier papers3
- DeMix: Debugging Training Data with Mixed Data Error Types by Investigating Influence VectorsJiale Deng, Yanyan Shen, Xiaogang Shi, Junjun ChaiKDD 2026
- DMCO: Budget-Aware Co-Optimization of Data Cleaning and AutoMLXiaoou Ding, Zekai Qian, Siying Chen, Hongbin Hu et al.ICML 2026
- Cleaning both Data Errors and Inaccurate Constraints on Numerical Sequential DataXiaoou Ding, Muyun Zhou, Yida Liu, Chen Wang et al.VLDB 2025
Builds on9
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan et al.VLDB 2021 · 484 citations
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang et al.ICDE 2021 · 127 citations
- Horizon: Scalable Dependency-driven Data CleaningEl Kindi Rezig, Mourad Ouzzani, Walid G. Aref, Ahmed K. Elmagarmid et al.VLDB 2021 · 95 citations
- Automatic Data Repair: Are We Ready to Deploy?Wei Ni, Xiaoye Miao, Xiangyu Zhao, Yangyang Wu et al.VLDB 2024 · 26 citations
- TSDDISCOVER: Discovering Data Dependency for Time Series DataXiaoou Ding, Yingze Li, Hongzhi Wang, Chen Wang et al.ICDE 2024 · 11 citations
Related papers
- GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language ModelsMengyi Yan, Yaoshu Wang, Yue Wang, Xiaoye Miao et al.SIGMOD 2025 · 12 citations
- MTSClean: Efficient Constraint-based Cleaning for Multi-Dimensional Time Series DataXiaoou Ding, Yichen Song, Hongzhi Wang, Chen Wang et al.VLDB 2024 · 9 citations
- SHoTClean: Bridging Soft and Hard Constraints for Multivariate Time Series CleaningZiquan Fang, Wei Shao, Zheqi Lu, Lu Chen et al.SIGMOD 2026
- SAGA: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning ApplicationsShafaq Siddiqi, Roman Kern, Matthias BoehmSIGMOD 2024 · 24 citations
- MINOR: Multivariate Time Series Iterative Cleaning AlgorithmAoqian Zhang, Yinru Sun, Pengxiang Hao, Yifeng Gong et al.ICDE 2026
