Deep Entity Matching with Pre-Trained Language Models
Yuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan, Wang-Chiew Tan
Abstract
We present Ditto, a novel entity matching system based on pre-trained Transformer-based language models. We fine-tune and cast EM as a sequence-pair classification problem to leverage such models with a simple architecture. Our experiments show that a straight-forward application of language models such as BERT, DistilBERT, or RoBERTa pre-trained on large text corpora already significantly improves the matching quality and outperforms previous state-of-the-art (SOTA), by up to 29% of F1 score on benchmark datasets. We also developed three optimization techniques to further improve Ditto's matching capability. Ditto allows domain knowledge to be injected by highlighting important pieces of input information that may be of interest when making matching decisions. Ditto also summarizes strings that are too long so that only the essential information is retained and used for EM. Finally, Ditto adapts a SOTA technique on data augmentation for text to EM to augment the training data with (difficult) examples. This way, Ditto is forced to learn "harder" to improve the model's matching capability. The optimizations we developed further boost the performance of Ditto by up to 9.8%. Perhaps more surprisingly, we establish that Ditto can achieve the previous SOTA results with at most half the number of labeled data. Finally, we demonstrate Ditto's effectiveness on a real-world large-scale EM task. On matching two company datasets consisting of 789K and 412K records, Ditto achieves a high F1 score of 96.5%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7106bc54-06c9-4dd6-a8e8-52cd1bb21d3fCited by top-tier papers97
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu et al.VLDB 2021 · 2,406 citations
- Can Foundation Models Wrangle Your Data?Avanika Narayan, Ines Chami, Laurel J. Orr, Christopher RéVLDB 2023 · 325 citations
- Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation LearningGrace Fan, Jin Wang, Yuliang Li, Dan Zhang et al.VLDB 2023 · 139 citations
- How Large Language Models Will Disrupt Data ManagementRaul Castro Fernandez, Aaron J. Elmore, Michael J. Franklin, Sanjay Krishnan et al.VLDB 2023 · 127 citations
- Deep Learning for Blocking in Entity Matching: A Design Space ExplorationSaravanan Thirumuruganathan, Han Li, Nan Tang, Mourad Ouzzani et al.VLDB 2021 · 109 citations
Builds on4
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- Snippext: Semi-supervised Opinion Mining with Augmented DataZhengjie Miao, Yuliang Li, Xiaolan Wang, Wang-Chiew TanWWW 2020 · 94 citations
- A Comprehensive Benchmark Framework for Active Learning Methods in Entity MatchingVenkata Vamsikrishna Meduri, Lucian Popa, Prithviraj Sen, Mohamed SarwatSIGMOD 2020 · 50 citations
- TaPas: Weakly Supervised Table Parsing via Pre-trainingJonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller, Francesco Piccinno et al.ACL 2020 · 19 citations
Related papers
- Analyzing How BERT Performs Entity MatchingMatteo Paganelli, Francesco Del Buono, Andrea Baraldi, Francesco GuerraVLDB 2022 · 35 citations
- BOND: BERT-Assisted Open-Domain Named Entity Recognition with Distant SupervisionChen Liang, Yue Yu, Haoming Jiang, Siawpeng Er et al.KDD 2020 · 118 citations
- Entity-aware Transformers for Entity SearchEmma J. Gerritse, Faegheh Hasibi, Arjen P. de VriesSIGIR 2022 · 27 citations
- Using Prior Knowledge to Guide BERT's Attention in Semantic Textual Matching TasksTingyu Xia, Yue Wang, Yuan Tian, Yi ChangWWW 2021 · 56 citations
- Pre-trained Embeddings for Entity Resolution: An Experimental AnalysisAlexandros Zeakis, George Papadakis, Dimitrios Skoutas, Manolis KoubarakisVLDB 2023 · 63 citations
