Pre-trained Embeddings for Entity Resolution: An Experimental Analysis
Alexandros Zeakis, George Papadakis, Dimitrios Skoutas, Manolis Koubarakis
Abstract
Many recent works on Entity Resolution (ER) leverage Deep Learning techniques involving language models to improve effectiveness. This is applied to both main steps of ER, i.e., blocking and matching. Several pre-trained embeddings have been tested, with the most popular ones being fastText and variants of the BERT model. However, there is no detailed analysis of their pros and cons. To cover this gap, we perform a thorough experimental analysis of 12 popular language models over 17 established benchmark datasets. First, we assess their vectorization overhead for converting all input entities into dense embeddings vectors. Second, we investigate their blocking performance, performing a detailed scalability analysis, and comparing them with the state-of-the-art deep learning-based blocking method. Third, we conclude with their relative performance for both supervised and unsupervised matching. Our experimental results provide novel insights into the strengths and weaknesses of the main language models, facilitating researchers and practitioners to select the most suitable ones in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 48eb8241-ecc8-46dd-b3c9-08e115a7bb9aCited by top-tier papers9
- GFM-RAG: Graph Foundation Model for Retrieval Augmented GenerationLinhao Luo, Zicheng Zhao, Reza Haffari, Dinh Phung et al.NeurIPS 2025 · 54 citations
- Watchog: A Light-weight Contrastive Learning based Framework for Column AnnotationZhengjie Miao, Jin WangSIGMOD 2024 · 14 citations
- Progressive Entity Matching: A Design Space ExplorationJakub Maciejewski, Konstantinos Nikoletos, George Papadakis, Yannis VelegrakisSIGMOD 2025 · 8 citations
- Learning from Natural Language Explanations for Generalizable Entity MatchingSomin Wadhwa, Adit Krishnan, Runhui Wang, Byron C. Wallace et al.EMNLP 2024 · 3 citations
- Can we trust LLM Self-Explanations for Entity Resolution?Tommaso Teofili, Donatella Firmani, Nick Koudas, Paolo Merialdo et al.VLDB 2026 · 2 citations
Builds on12
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan et al.VLDB 2021 · 484 citations
- Large Dual Encoders Are Generalizable RetrieversJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai et al.EMNLP 2022 · 145 citations
Related papers
- Improving the Efficiency and Effectiveness for BERT-based Entity ResolutionBing Li, Yukai Miao, Yaoshu Wang, Yifang Sun et al.AAAI 2021 · 45 citations
- Deep Learning for Blocking in Entity Matching: A Design Space ExplorationSaravanan Thirumuruganathan, Han Li, Nan Tang, Mourad Ouzzani et al.VLDB 2021 · 109 citations
- A Critical Re-evaluation of Record Linkage Benchmarks for Learning-Based Matching AlgorithmsGeorge Papadakis, Nishadi Kirielle, Peter Christen, Themis PalpanasICDE 2024 · 8 citations
- Deep Indexed Active Learning for Matching Heterogeneous Entity RepresentationsArjit Jain, Sunita Sarawagi, Prithviraj SenVLDB 2022 · 32 citations
- Analyzing How BERT Performs Entity MatchingMatteo Paganelli, Francesco Del Buono, Andrea Baraldi, Francesco GuerraVLDB 2022 · 35 citations
