MDedup: Duplicate Detection with Matching Dependencies
Ioannis K. Koumarelas, Thorsten Papenbrock, Felix Naumann
摘要
Duplicate detection is an integral part of data cleaning and serves to identify multiple representations of same real-world entities in (relational) datasets. Existing duplicate detection approaches are effective, but they are also hard to parameterize or require a lot of pre-labeled training data. Both parameterization and pre-labeling are at least domain-specific if not dataset-specific, which is a problem if a new dataset needs to be cleaned. For this reason, we propose a novel, rule-based and fully automatic duplicate detection approach that is based on matching dependencies (MDs). Our system uses automatically discovered MDs, various dataset features, and known gold standards to train a model that selects MDs as duplicate detection rules. Once trained, the model can select useful MDs for duplicate detection on any new dataset. To increase the generally low recall of MD-based data cleaning approaches, we propose an additional boosting step. Our experiments show that this approach reaches up to 94% F-measure and 100% precision on our evaluation datasets, which are good numbers considering that the system does not require domain or target data-specific configuration.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and BeyondZhengjie Miao, Yuliang Li, Xiaolan WangSIGMOD 2021 · 被引用 63 次
- Parallel Rule Discovery from Large Datasets by SamplingWenfei Fan, Ziyan Han, Yaoshu Wang, Min XieSIGMOD 2022 · 被引用 21 次
- Learning Over Dirty Data Without CleaningJose Picado, John Davis, Arash Termehchy, Ga Young LeeSIGMOD 2020 · 被引用 13 次
- Discovering Top-k Rules using Subjective and Objective CriteriaWenfei Fan, Ziyan Han, Yaoshu Wang, Min XieSIGMOD 2023 · 被引用 10 次
- Deep and Collective Entity Resolution in ParallelTing Deng, Wenfei Fan, Ping Lu, Xiaomeng Luo 等ICDE 2022 · 被引用 6 次
相关 Paper
- A Comprehensive Benchmark Framework for Active Learning Methods in Entity MatchingVenkata Vamsikrishna Meduri, Lucian Popa, Prithviraj Sen, Mohamed SarwatSIGMOD 2020 · 被引用 50 次
- Sparcle: Boosting the Accuracy of Data Cleaning Systems through Spatial AwarenessYuchuan Huang, Mohamed F. MokbelVLDB 2024 · 被引用 3 次
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan 等VLDB 2021 · 被引用 484 次
- Efficient Differential Dependency DiscoveryShulei Kuang, Honghui Yang, Zijing Tan, Shuai MaVLDB 2024 · 被引用 4 次
- Parallel Discrepancy Detection and Incremental DetectionWenfei Fan, Chao Tian, Yanghao Wang, Qiang YinVLDB 2021 · 被引用 29 次
