Analyzing How BERT Performs Entity Matching
Matteo Paganelli, Francesco Del Buono, Andrea Baraldi, Francesco Guerra
Abstract
State-of-the-art Entity Matching (EM) approaches rely on transformer architectures, such as BERT , for generating highly contex-tualized embeddings of terms. The embeddings are then used to predict whether pairs of entity descriptions refer to the same real-world entity. BERT-based EM models demonstrated to be effective, but act as black-boxes for the users, who have limited insight into the motivations behind their decisions.
In this paper, we perform a multi-facet analysis of the components of pre-trained and fine-tuned BERT architectures applied to an EM task. The main findings resulting from our extensive experimental evaluation are (1) the fine-tuning process applied to the EM task mainly modifies the last layers of the BERT components, but in a different way on tokens belonging to descriptions of matching / non-matching entities; (2) the special structure of the EM datasets, where records are pairs of entity descriptions is recognized by BERT; (3) the pair-wise semantic similarity of tokens is not a key knowledge exploited by BERT-based EM models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 96652efb-1d09-45c3-91c5-235bc4587359Cited by top-tier papers8
- Pre-trained Embeddings for Entity Resolution: An Experimental AnalysisAlexandros Zeakis, George Papadakis, Dimitrios Skoutas, Manolis KoubarakisVLDB 2023 · 63 citations
- Automatic Data Repair: Are We Ready to Deploy?Wei Ni, Xiaoye Miao, Xiangyu Zhao, Yangyang Wu et al.VLDB 2024 · 26 citations
- Deep Active Alignment of Knowledge Graph Entities and SchemataJiacheng Huang, Zequn Sun, Qijin Chen, Xiaozhou Xu et al.SIGMOD 2023 · 10 citations
- A Critical Re-evaluation of Record Linkage Benchmarks for Learning-Based Matching AlgorithmsGeorge Papadakis, Nishadi Kirielle, Peter Christen, Themis PalpanasICDE 2024 · 8 citations
- BEACON: Budget-Aware Entity Matching Across DomainsNicholas Pulsone, Roee Shraga, Gregory GorenSIGMOD 2026 · 2 citations
Builds on4
- On Identifiability in TransformersGino Brunner, Yang Liu, Damian Pascual, Oliver Richter et al.ICLR 2020 · 210 citations
- Perturbed Masking: Parameter-free Probing for Analyzing and Interpreting BERTZhiyong Wu, Yun Chen, Ben Kao, Qun LiuACL 2020 · 158 citations
- Deep Learning for Blocking in Entity Matching: A Design Space ExplorationSaravanan Thirumuruganathan, Han Li, Nan Tang, Mourad Ouzzani et al.VLDB 2021 · 109 citations
- Dual-Objective Fine-Tuning of BERT for Entity MatchingRalph Peeters, Christian BizerVLDB 2021 · 71 citations
Related papers
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan et al.VLDB 2021 · 484 citations
- Improving the Efficiency and Effectiveness for BERT-based Entity ResolutionBing Li, Yukai Miao, Yaoshu Wang, Yifang Sun et al.AAAI 2021 · 45 citations
- Improving Entity Linking by Modeling Latent Entity Type InformationShuang Chen, Jinpeng Wang, Feng Jiang, Chin-Yew LinAAAI 2020 · 71 citations
- Using Prior Knowledge to Guide BERT's Attention in Semantic Textual Matching TasksTingyu Xia, Yue Wang, Yuan Tian, Yi ChangWWW 2021 · 56 citations
- Entity-aware Transformers for Entity SearchEmma J. Gerritse, Faegheh Hasibi, Arjen P. de VriesSIGIR 2022 · 27 citations
