Ground Truth Inference for Weakly Supervised Entity Matching
Renzhi Wu, Alexander Bendeck, Xu Chu, Yeye He
Abstract
Entity matching (EM) refers to the problem of identifying pairs of data records in one or more relational tables that refer to the same entity in the real world. Supervised machine learning (ML) models currently achieve state-of-the-art matching performance; however, they require a large number of labeled examples, which are often expensive or infeasible to obtain. This has inspired us to approach data labeling for EM using weak supervision. In particular, we use the labeling function abstraction popularized by Snorkel, where each labeling function (LF) is a user-provided program that can generate many noisy match/non-match labels quickly and cheaply. Given a set of user-written LFs, the quality of data labeling depends on a labeling model to accurately infer the ground-truth labels. In this work, we first propose a simple but powerful labeling model for general weak supervision tasks. Then, we tailor the labeling model specifically to the task of entity matching by considering the EM-specific transitivity property. The general form of our labeling model is simple while substantially outperforming the best existing method across ten general weak supervision datasets. To tailor the labeling model for EM, we formulate an approach to ensure that the final predictions of the labeling model satisfy the transitivity property required in EM, utilizing an exact solution where possible and an ML-based approximation in remaining cases. On two single-table and nine two-table real-world EM datasets, we show that our labeling model results in a 9% higher F1 score on average than the best existing method. We also show that a deep learning EM end model (DeepMatcher) trained on labels generated from our weak supervision approach is comparable to an end model trained using tens of thousands of ground-truth labels, demonstrating that our approach can significantly reduce the labeling efforts required in EM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 53abbcec-7c8d-481b-9568-703e862b7d90Cited by top-tier papers2
- Blocker and Matcher Can Mutually Benefit: A Co-Learning Framework for Low-Resource Entity ResolutionShiwen Wu, Qiyu Wu, Honghua Dong, Wen Hua et al.VLDB 2024 · 10 citations
- Adaptive Graph Refinement and Label Propagation with LLMs for Cost-Effective Entity ResolutionHongtao Wang, Renchi Yang, Haoran Zheng, Xiangyu KeKDD 2026 · 1 citation
Builds on13
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan et al.VLDB 2021 · 484 citations
- Fast and Three-rious: Speeding Up Weak Supervision with Triplet MethodsDaniel Y. Fu, Mayee F. Chen, Frederic Sala, Sarah M. Hooper et al.ICML 2020 · 130 citations
- ZeroER: Entity Resolution using Zero Labeled ExamplesRenzhi Wu, Sanya Chaba, Saurabh Sawlani, Xu Chu et al.SIGMOD 2020 · 77 citations
- End-to-End Weak SupervisionSalva Rühling Cachay, Benedikt Boecking, Artur DubrawskiNeurIPS 2021 · 48 citations
- Strength from Weakness: Fast Learning Using Weak SupervisionJoshua Robinson, Stefanie Jegelka, Suvrit SraICML 2020 · 35 citations
Related papers
- A Comprehensive Benchmark Framework for Active Learning Methods in Entity MatchingVenkata Vamsikrishna Meduri, Lucian Popa, Prithviraj Sen, Mohamed SarwatSIGMOD 2020 · 50 citations
- Deep Learning for Blocking in Entity Matching: A Design Space ExplorationSaravanan Thirumuruganathan, Han Li, Nan Tang, Mourad Ouzzani et al.VLDB 2021 · 109 citations
- Automating Entity Matching Model DevelopmentPei Wang, Weiling Zheng, Jiannan Wang, Jian PeiICDE 2021 · 13 citations
- BEACON: Budget-Aware Entity Matching Across DomainsNicholas Pulsone, Roee Shraga, Gregory GorenSIGMOD 2026 · 2 citations
- Learning from Natural Language Explanations for Generalizable Entity MatchingSomin Wadhwa, Adit Krishnan, Runhui Wang, Byron C. Wallace et al.EMNLP 2024 · 3 citations
