MDedup: Duplicate Detection with Matching Dependencies
Ioannis K. Koumarelas, Thorsten Papenbrock, Felix Naumann
Abstract
Duplicate detection is an integral part of data cleaning and serves to identify multiple representations of same real-world entities in (relational) datasets. Existing duplicate detection approaches are effective, but they are also hard to parameterize or require a lot of pre-labeled training data. Both parameterization and pre-labeling are at least domain-specific if not dataset-specific, which is a problem if a new dataset needs to be cleaned. For this reason, we propose a novel, rule-based and fully automatic duplicate detection approach that is based on matching dependencies (MDs). Our system uses automatically discovered MDs, various dataset features, and known gold standards to train a model that selects MDs as duplicate detection rules. Once trained, the model can select useful MDs for duplicate detection on any new dataset. To increase the generally low recall of MD-based data cleaning approaches, we propose an additional boosting step. Our experiments show that this approach reaches up to 94% F-measure and 100% precision on our evaluation datasets, which are good numbers considering that the system does not require domain or target data-specific configuration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c8155770-b08d-4e39-9a36-cd73ce06b83bCited by top-tier papers11
- Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and BeyondZhengjie Miao, Yuliang Li, Xiaolan WangSIGMOD 2021 · 63 citations
- Parallel Rule Discovery from Large Datasets by SamplingWenfei Fan, Ziyan Han, Yaoshu Wang, Min XieSIGMOD 2022 · 21 citations
- Learning Over Dirty Data Without CleaningJose Picado, John Davis, Arash Termehchy, Ga Young LeeSIGMOD 2020 · 13 citations
- Discovering Top-k Rules using Subjective and Objective CriteriaWenfei Fan, Ziyan Han, Yaoshu Wang, Min XieSIGMOD 2023 · 10 citations
- Deep and Collective Entity Resolution in ParallelTing Deng, Wenfei Fan, Ping Lu, Xiaomeng Luo et al.ICDE 2022 · 6 citations
Related papers
- A Comprehensive Benchmark Framework for Active Learning Methods in Entity MatchingVenkata Vamsikrishna Meduri, Lucian Popa, Prithviraj Sen, Mohamed SarwatSIGMOD 2020 · 50 citations
- Sparcle: Boosting the Accuracy of Data Cleaning Systems through Spatial AwarenessYuchuan Huang, Mohamed F. MokbelVLDB 2024 · 3 citations
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan et al.VLDB 2021 · 484 citations
- Efficient Differential Dependency DiscoveryShulei Kuang, Honghui Yang, Zijing Tan, Shuai MaVLDB 2024 · 4 citations
- Parallel Discrepancy Detection and Incremental DetectionWenfei Fan, Chao Tian, Yanghao Wang, Qiang YinVLDB 2021 · 29 citations
