Schema Matching using Pre-Trained Language Models
Yunjia Zhang, Avrilia Floratou, Joyce Cahoon, Subru Krishnan, Andreas C. Müller, Dalitso Banda, Fotis Psallidas, Jignesh M. Patel
Abstract
Schema matching over relational data has been studied for more than two decades. However, the state-of-the-art methods do not address key modern-day challenges encountered in real customer scenarios, namely: 1) no access to the source (customer) data due to privacy constraints, 2) target schema with a much larger number of entities and attributes compared to the source schema, and 3) different but semantically equivalent entity and attribute names in the source and target schemata. In this paper, we address these shortcomings. Using real-world customer schemata, we demonstrate that existing linguistic matching approaches have low accuracy. Next, we propose the Learned Schema Mapper (LSM), a novel linguistic schema matching system that leverages the natural language understanding capabilities of pre-trained language models to improve the overall accuracy. Combining this with active learning and a smart attribute selection strategy that selects the most informative attributes for users to label, LSM can significantly reduce the overall human labeling cost. Experimental results demonstrate that users can correctly match their full schema while saving as much as 81% of the labeling cost compared to manual labeling.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 7dc1347e-540c-4252-b1a9-9aa351c6cf70Cited by top-tier papers7
- GPTuner: A Manual-Reading Database Tuning System via GPT-Guided Bayesian OptimizationJiale Lao, Yibo Wang, Yufei Li, Jianping Wang et al.VLDB 2024 · 76 citations
- Magneto: Combining Small and Large Language Models for Schema MatchingYurong Liu, Eduardo H. M. Pena, Aécio S. R. Santos, Eden Wu et al.VLDB 2025 · 32 citations
- Sphinteract: Resolving Ambiguities in NL2SQL Through User InteractionFuheng Zhao, Shaleen Deep, Fotis Psallidas, Avrilia Floratou et al.VLDB 2025 · 12 citations
- OmniMatch: Joinability Discovery in Data ProductsChristos Koutras, Jiani Zhang, Xiao Qin, Chuan Lei et al.VLDB 2025 · 3 citations
- BDIViz: An Interactive Visualization System for Biomedical Schema Matching with LLM-Powered ValidationEden Wu, Dishita G. Turakhia, Guande Wu, Christos Koutras et al.IEEE VIS 2025 · 3 citations
Related papers
- In Situ Neural Relational Schema MatcherXingyu Du, Gongsheng Yuan, Sai Wu, Gang Chen et al.ICDE 2024 · 4 citations
- Bootstrapping Self-Improvement of Language Model Programs for Zero-Shot Schema MatchingNabeel Seedat, Mihaela van der SchaarICML 2025
- The Battleship Approach to the Low Resource Entity Matching ProblemBar Genossar, Avigdor Gal, Roee ShragaSIGMOD 2024 · 6 citations
- Deep Active Alignment of Knowledge Graph Entities and SchemataJiacheng Huang, Zequn Sun, Qijin Chen, Xiaozhou Xu et al.SIGMOD 2023 · 10 citations
- BEACON: Budget-Aware Entity Matching Across DomainsNicholas Pulsone, Roee Shraga, Gregory GorenSIGMOD 2026 · 2 citations
