Dual-Objective Fine-Tuning of BERT for Entity Matching
Ralph Peeters, Christian Bizer
摘要
An increasing number of data providers have adopted shared numbering schemes such as GTIN, ISBN, DUNS, or ORCID numbers for identifying entities in the respective domain. This means for data integration that shared identifiers are often available for a subset of the entity descriptions to be integrated while such identifiers are not available for others. The challenge in these settings is to learn a matcher for entity descriptions without identifiers using the entity descriptions containing identifiers as training data. The task can be approached by learning a binary classifier which distinguishes pairs of entity descriptions for the same real-world entity from descriptions of different entities. The task can also be modeled as a multi-class classification problem by learning classifiers for identifying descriptions of individual entities. We present a dual-objective training method for BERT, called JointBERT, which combines binary matching and multi-class classification, forcing the model to predict the entity identifier for each entity description in a training pair in addition to the match/non-match decision. Our evaluation across five entity matching benchmark datasets shows that dual-objective training can increase the matching performance for seen products by 1% to 5% F1 compared to single-objective Transformer-based methods, given that enough training data is available for both objectives. In order to gain a deeper understanding of the strengths and weaknesses of the proposed method, we compare JointBERT to several other BERT-based matching methods as well as baseline systems along a set of specific matching challenges. This evaluation shows that JointBERT, given enough training data for both objectives, outperforms the other methods on tasks involving seen products, while it underperforms for unseen products. Using a combination of LIME explanations and domain-specific word classes, we analyze the matching decisions of the different deep learning models and conclude that BERT-based models are better at focusing on relevant word classes compared to RNN-based models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Pre-trained Embeddings for Entity Resolution: An Experimental AnalysisAlexandros Zeakis, George Papadakis, Dimitrios Skoutas, Manolis KoubarakisVLDB 2023 · 被引用 63 次
- Cost-Effective In-Context Learning for Entity Resolution: A Design Space ExplorationMeihao Fan, Xiaoyue Han, Ju Fan, Chengliang Chai 等ICDE 2024 · 被引用 40 次
- Analyzing How BERT Performs Entity MatchingMatteo Paganelli, Francesco Del Buono, Andrea Baraldi, Francesco GuerraVLDB 2022 · 被引用 35 次
- Sudowoodo: Contrastive Self-supervised Learning for Multi-purpose Data Integration and PreparationRunhui Wang, Yuliang Li, Jin WangICDE 2023 · 被引用 32 次
- Geospatial Entity ResolutionPasquale Balsebre, Dezhong Yao, Gao Cong, Zhen HaiWWW 2022 · 被引用 21 次
它引用的顶会 Paper5
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralisationJunjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig 等ICML 2020 · 被引用 1,132 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan 等VLDB 2021 · 被引用 484 次
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 被引用 417 次
相关 Paper
- A Critical Re-evaluation of Record Linkage Benchmarks for Learning-Based Matching AlgorithmsGeorge Papadakis, Nishadi Kirielle, Peter Christen, Themis PalpanasICDE 2024 · 被引用 8 次
- The Battleship Approach to the Low Resource Entity Matching ProblemBar Genossar, Avigdor Gal, Roee ShragaSIGMOD 2024 · 被引用 6 次
- Blocker and Matcher Can Mutually Benefit: A Co-Learning Framework for Low-Resource Entity ResolutionShiwen Wu, Qiyu Wu, Honghua Dong, Wen Hua 等VLDB 2024 · 被引用 10 次
- Learning from Natural Language Explanations for Generalizable Entity MatchingSomin Wadhwa, Adit Krishnan, Runhui Wang, Byron C. Wallace 等EMNLP 2024 · 被引用 3 次
- BEACON: Budget-Aware Entity Matching Across DomainsNicholas Pulsone, Roee Shraga, Gregory GorenSIGMOD 2026 · 被引用 2 次
