Horizon: Scalable Dependency-driven Data Cleaning
El Kindi Rezig, Mourad Ouzzani, Walid G. Aref, Ahmed K. Elmagarmid, Ahmed R. Mahmood, Michael Stonebraker
Abstract
A large class of data repair algorithms rely on integrity constraints to detect and repair errors. A well-studied class of constraints is Functional Dependencies (FDs, for short). Although there has been an increased interest in developing general data cleaning systems for a myriad of data errors, scalability has been left behind. This is because current systems assume data cleaning is performed offline and in one iteration. However, developing data science pipelines is highly iterative and requires efficient cleaning techniques to scale to millions of records in seconds/minutes, not days. In our efforts to re-think the data cleaning stack and bring it to the era of data science, we introduce Horizon , an end-to-end FD repair system to address two key challenges: (1) Accuracy: Most existing FD repair techniques aim to produce repairs that minimize changes to the data that may lead to incorrect combinations of attribute values (or patterns). Horizon leverages the interaction between the data patterns induced by the various FDs, and subsequently selects repairs that preserve the most frequent patterns found in the original data, and hence leading to a better repair accuracy. (2) Scalability: Existing data cleaning systems struggle when dealing with large-scale real-world datasets. Horizon features a linear-time repair algorithm that scales to millions of records, and is orders-of-magnitude faster than state-of-the-art cleaning algorithms. A benchmark of Horizon against state-of-the-art cleaning systems on multiple datasets and metrics shows that Horizon consistently outperforms existing techniques in repair quality and scalability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f082b12-1a75-4c57-86e0-be45bfb939eeCited by top-tier papers20
- Automatic Data Repair: Are We Ready to Deploy?Wei Ni, Xiaoye Miao, Xiangyu Zhao, Yangyang Wu et al.VLDB 2024 · 26 citations
- Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial NetworksJinfeng Peng, Derong Shen, Nan Tang, Tieying Liu et al.VLDB 2023 · 24 citations
- DataPrism: Exposing Disconnect between Data and SystemsSainyam Galhotra, Anna Fariha, Raoni Lourenço, Juliana Freire et al.SIGMOD 2022 · 9 citations
- KGLiDS: A Platform for Semantic Abstraction, Linking, and Automation of Data ScienceMossad Helali, Niki Monjazeb, Shubham Vashisth, Philippe Carrier et al.ICDE 2024 · 7 citations
- Naive Bayes Classifiers over Missing Data: Decision and PoisoningSong Bian, Xiating Ouyang, Zhiwei Fan, Paraschos KoutrisICML 2024 · 5 citations
Builds on1
Related papers
- Repairing Entities using Star Constraints in Multirelational GraphsPeng Lin, Qi Song, Yinghui Wu, Jiaxing PiICDE 2020 · 7 citations
- Pattern Functional Dependencies for Data CleaningAbdulhakim Ali Qahtan, Nan Tang, Mourad Ouzzani, Yang Cao et al.VLDB 2020 · 42 citations
- Multivariate Time Series Cleaning under Speed ConstraintsAoqian Zhang, Zexue Wu, Yifeng Gong, Ye Yuan et al.SIGMOD 2025 · 4 citations
- Time Series Data Cleaning Under Expressive Constraints on Both Rows and ColumnsXiaoou Ding, Genglong Li, Hongzhi Wang, Chen Wang et al.ICDE 2024 · 7 citations
- Discovering Functional Dependencies through Hitting Set EnumerationTobias Bleifuß, Thorsten Papenbrock, Thomas Bläsius, Martin Schirneck et al.SIGMOD 2024 · 9 citations
