ADnEV: Cross-Domain Schema Matching using Deep Similarity Matrix Adjustment and Evaluation
Roee Shraga, Avigdor Gal, Haggai Roitman
Abstract
Schema matching is a process that serves in integrating structured and semi-structured data. Being a handy tool in multiple contemporary business and commerce applications, it has been investigated in the fields of databases, AI, Semantic Web, and data mining for many years. The core challenge still remains the ability to create quality algorithmic matchers, automatic tools for identifying correspondences among data concepts (e.g., database attributes). In this work, we offer a novel post processing step to schema matching that improves the final matching outcome without human intervention. We present a new mechanism, similarity matrix adjustment, to calibrate a matching result and propose an algorithm (dubbed ADnEV) that manipulates, using deep neural networks, similarity matrices, created by state-of-the-art algorithmic matchers. ADnEV learns two models that iteratively adjust and evaluate the original similarity matrix. We empirically demonstrate the effectiveness of the proposed algorithmic solution for improving matching results, using real-world benchmark ontology and schema sets. We show that ADnEV can generalize into new domains without the need to learn the domain terminology, thus allowing cross-domain learning. We also show ADnEV to be a powerful tool in handling schemata which matching is particularly challenging. Finally, we show the benefit of using ADnEV in a related integration task of ontology alignment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b2da031-9d40-466e-958f-774db5b4f28bCited by top-tier papers12
- Integrating Data Lake TablesAamod Khatiwada, Roee Shraga, Wolfgang Gatterbauer, Renée J. MillerVLDB 2023 · 59 citations
- Magneto: Combining Small and Large Language Models for Schema MatchingYurong Liu, Eduardo H. M. Pena, Aécio S. R. Santos, Eden Wu et al.VLDB 2025 · 32 citations
- Explaining Dataset Changes for Semantic Data Versioning with Explain-Da-VRoee Shraga, Renée J. MillerVLDB 2023 · 18 citations
- FlexER: Flexible Entity Resolution for Multiple IntentsBar Genossar, Roee Shraga, Avigdor GalSIGMOD 2023 · 15 citations
- Jellyfish: Instruction-Tuning Local Large Language Models for Data PreprocessingHaochen Zhang, Yuyang Dong, Chuan Xiao, Masafumi OyamadaEMNLP 2024 · 11 citations
Related papers
- Deep Active Alignment of Knowledge Graph Entities and SchemataJiacheng Huang, Zequn Sun, Qijin Chen, Xiaozhou Xu et al.SIGMOD 2023 · 10 citations
- Bootstrapping Self-Improvement of Language Model Programs for Zero-Shot Schema MatchingNabeel Seedat, Mihaela van der SchaarICML 2025
- In Situ Neural Relational Schema MatcherXingyu Du, Gongsheng Yuan, Sai Wu, Gang Chen et al.ICDE 2024 · 4 citations
- Domain Adaptation for Deep Entity ResolutionJianhong Tu, Ju Fan, Nan Tang, Peng Wang et al.SIGMOD 2022 · 46 citations
- Alignment Attention by Matching Key and Query DistributionsShujian Zhang, Xinjie Fan, Huangjie Zheng, Korawat Tanwisuth et al.NeurIPS 2021 · 20 citations
