Learning Over Dirty Data Without Cleaning
Jose Picado, John Davis, Arash Termehchy, Ga Young Lee
摘要
Real-world datasets are dirty and contain many errors. Examples of these issues are violations of integrity constraints, duplicates, and inconsistencies in representing data values and entities. Learning over dirty databases may result in inaccurate models. Users have to spend a great deal of time and effort to repair data errors and create a clean database for learning. Moreover, as the information required to repair these errors is not often available, there may be numerous possible clean versions for a dirty database. We propose DLearn, a novel relational learning system that learns directly over dirty databases effectively and efficiently without any preprocessing. DLearn leverages database constraints to learn accurate relational models over inconsistent and heterogeneous data. Its learned models represent patterns over all possible clean instances of the data in a usable form. Our empirical study indicates that DLearn learns accurate models over large real-world databases efficiently.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Leva: Boosting Machine Learning Performance with Relational Embedding Data AugmentationZixuan Zhao, Raul Castro FernandezSIGMOD 2022 · 被引用 19 次
- Naive Bayes Classifiers over Missing Data: Decision and PoisoningSong Bian, Xiating Ouyang, Zhiwei Fan, Paraschos KoutrisICML 2024 · 被引用 5 次
- In-Database Data ImputationMassimo Perini, Milos NikolicSIGMOD 2024 · 被引用 5 次
- Certain and Approximately Certain Models for Statistical LearningCheng Zhen, Nischal Aryal, Arash Termehchy, Amandeep Singh ChabadaSIGMOD 2024 · 被引用 4 次
它引用的顶会 Paper1
相关 Paper
- Cleaning Time Series under Seasonal and Trend ConstraintsZijie Chen, Aoqian Zhang, Shaoxu SongSIGMOD 2026
- PrIU: A Provenance-Based Approach for Incrementally Updating Regression ModelsYinjun Wu, Val Tannen, Susan B. DavidsonSIGMOD 2020 · 被引用 26 次
- Outliers: The Good, the Bad and the UglyShenglin Chen, Wenfei Fan, Ruochun JinSIGMOD 2026
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification TasksPeng Li, Xi Rao, Jennifer Blase, Yue Zhang 等ICDE 2021 · 被引用 127 次
- Generalizable Data Cleaning of Tabular Data in Latent SpaceEduardo Souza dos Reis, Mohamed Abdelaal, Carsten BinnigVLDB 2024 · 被引用 3 次
