Horizon: Scalable Dependency-driven Data Cleaning
El Kindi Rezig, Mourad Ouzzani, Walid G. Aref, Ahmed K. Elmagarmid, Ahmed R. Mahmood, Michael Stonebraker
摘要
A large class of data repair algorithms rely on integrity constraints to detect and repair errors. A well-studied class of constraints is Functional Dependencies (FDs, for short). Although there has been an increased interest in developing general data cleaning systems for a myriad of data errors, scalability has been left behind. This is because current systems assume data cleaning is performed offline and in one iteration. However, developing data science pipelines is highly iterative and requires efficient cleaning techniques to scale to millions of records in seconds/minutes, not days. In our efforts to re-think the data cleaning stack and bring it to the era of data science, we introduce Horizon , an end-to-end FD repair system to address two key challenges: (1) Accuracy: Most existing FD repair techniques aim to produce repairs that minimize changes to the data that may lead to incorrect combinations of attribute values (or patterns). Horizon leverages the interaction between the data patterns induced by the various FDs, and subsequently selects repairs that preserve the most frequent patterns found in the original data, and hence leading to a better repair accuracy. (2) Scalability: Existing data cleaning systems struggle when dealing with large-scale real-world datasets. Horizon features a linear-time repair algorithm that scales to millions of records, and is orders-of-magnitude faster than state-of-the-art cleaning algorithms. A benchmark of Horizon against state-of-the-art cleaning systems on multiple datasets and metrics shows that Horizon consistently outperforms existing techniques in repair quality and scalability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Automatic Data Repair: Are We Ready to Deploy?Wei Ni, Xiaoye Miao, Xiangyu Zhao, Yangyang Wu 等VLDB 2024 · 被引用 26 次
- Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial NetworksJinfeng Peng, Derong Shen, Nan Tang, Tieying Liu 等VLDB 2023 · 被引用 24 次
- DataPrism: Exposing Disconnect between Data and SystemsSainyam Galhotra, Anna Fariha, Raoni Lourenço, Juliana Freire 等SIGMOD 2022 · 被引用 9 次
- KGLiDS: A Platform for Semantic Abstraction, Linking, and Automation of Data ScienceMossad Helali, Niki Monjazeb, Shubham Vashisth, Philippe Carrier 等ICDE 2024 · 被引用 7 次
- Naive Bayes Classifiers over Missing Data: Decision and PoisoningSong Bian, Xiating Ouyang, Zhiwei Fan, Paraschos KoutrisICML 2024 · 被引用 5 次
它引用的顶会 Paper1
相关 Paper
- Repairing Entities using Star Constraints in Multirelational GraphsPeng Lin, Qi Song, Yinghui Wu, Jiaxing PiICDE 2020 · 被引用 7 次
- Pattern Functional Dependencies for Data CleaningAbdulhakim Ali Qahtan, Nan Tang, Mourad Ouzzani, Yang Cao 等VLDB 2020 · 被引用 42 次
- Multivariate Time Series Cleaning under Speed ConstraintsAoqian Zhang, Zexue Wu, Yifeng Gong, Ye Yuan 等SIGMOD 2025 · 被引用 4 次
- Time Series Data Cleaning Under Expressive Constraints on Both Rows and ColumnsXiaoou Ding, Genglong Li, Hongzhi Wang, Chen Wang 等ICDE 2024 · 被引用 7 次
- Discovering Functional Dependencies through Hitting Set EnumerationTobias Bleifuß, Thorsten Papenbrock, Thomas Bläsius, Martin Schirneck 等SIGMOD 2024 · 被引用 9 次
