Ground Truth Inference for Weakly Supervised Entity Matching
Renzhi Wu, Alexander Bendeck, Xu Chu, Yeye He
摘要
Entity matching (EM) refers to the problem of identifying pairs of data records in one or more relational tables that refer to the same entity in the real world. Supervised machine learning (ML) models currently achieve state-of-the-art matching performance; however, they require a large number of labeled examples, which are often expensive or infeasible to obtain. This has inspired us to approach data labeling for EM using weak supervision. In particular, we use the labeling function abstraction popularized by Snorkel, where each labeling function (LF) is a user-provided program that can generate many noisy match/non-match labels quickly and cheaply. Given a set of user-written LFs, the quality of data labeling depends on a labeling model to accurately infer the ground-truth labels. In this work, we first propose a simple but powerful labeling model for general weak supervision tasks. Then, we tailor the labeling model specifically to the task of entity matching by considering the EM-specific transitivity property. The general form of our labeling model is simple while substantially outperforming the best existing method across ten general weak supervision datasets. To tailor the labeling model for EM, we formulate an approach to ensure that the final predictions of the labeling model satisfy the transitivity property required in EM, utilizing an exact solution where possible and an ML-based approximation in remaining cases. On two single-table and nine two-table real-world EM datasets, we show that our labeling model results in a 9% higher F1 score on average than the best existing method. We also show that a deep learning EM end model (DeepMatcher) trained on labels generated from our weak supervision approach is comparable to an end model trained using tens of thousands of ground-truth labels, demonstrating that our approach can significantly reduce the labeling efforts required in EM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Blocker and Matcher Can Mutually Benefit: A Co-Learning Framework for Low-Resource Entity ResolutionShiwen Wu, Qiyu Wu, Honghua Dong, Wen Hua 等VLDB 2024 · 被引用 10 次
- Adaptive Graph Refinement and Label Propagation with LLMs for Cost-Effective Entity ResolutionHongtao Wang, Renchi Yang, Haoran Zheng, Xiangyu KeKDD 2026 · 被引用 1 次
它引用的顶会 Paper13
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan 等VLDB 2021 · 被引用 484 次
- Fast and Three-rious: Speeding Up Weak Supervision with Triplet MethodsDaniel Y. Fu, Mayee F. Chen, Frederic Sala, Sarah M. Hooper 等ICML 2020 · 被引用 130 次
- ZeroER: Entity Resolution using Zero Labeled ExamplesRenzhi Wu, Sanya Chaba, Saurabh Sawlani, Xu Chu 等SIGMOD 2020 · 被引用 77 次
- End-to-End Weak SupervisionSalva Rühling Cachay, Benedikt Boecking, Artur DubrawskiNeurIPS 2021 · 被引用 48 次
- Strength from Weakness: Fast Learning Using Weak SupervisionJoshua Robinson, Stefanie Jegelka, Suvrit SraICML 2020 · 被引用 35 次
相关 Paper
- A Comprehensive Benchmark Framework for Active Learning Methods in Entity MatchingVenkata Vamsikrishna Meduri, Lucian Popa, Prithviraj Sen, Mohamed SarwatSIGMOD 2020 · 被引用 50 次
- Deep Learning for Blocking in Entity Matching: A Design Space ExplorationSaravanan Thirumuruganathan, Han Li, Nan Tang, Mourad Ouzzani 等VLDB 2021 · 被引用 109 次
- Automating Entity Matching Model DevelopmentPei Wang, Weiling Zheng, Jiannan Wang, Jian PeiICDE 2021 · 被引用 13 次
- BEACON: Budget-Aware Entity Matching Across DomainsNicholas Pulsone, Roee Shraga, Gregory GorenSIGMOD 2026 · 被引用 2 次
- Learning from Natural Language Explanations for Generalizable Entity MatchingSomin Wadhwa, Adit Krishnan, Runhui Wang, Byron C. Wallace 等EMNLP 2024 · 被引用 3 次
