SemStruct: Contextualizing Semantic Embeddings with Structural Information for Schema Matching
Inwon Kang, Kavitha Srinivas, Nandana Mihindukulasooriya, Sola Shirai, Parikshit Ram, Horst Samulowitz, Oshani Seneviratne
Abstract
Schema matching is a fundamental step in integrating heterogeneous data sources. While Pre-trained Language Models (PLMs) have revolutionized this task by capturing linguistic semantics, they typically process tabular data as serialized text sequences of standalone column descriptions. This serialization discards critical structural information -- specifically, the row-level co-occurrences, i.e. the relational context -- forcing models to rely solely on column header semantics or standalone distributions. To bridge this gap, we propose SemStruct, a framework that joins the semantic power of frozen PLMs with the structural inductive bias of Graph Neural Networks (GNNs). We model the table as a heterogeneous graph where columns and values are nodes connected by rows, allowing the GNN to propagate disambiguating context across the structure. Unlike other state-of-the-art methods that require proprietary LLM access and fine-tuning of language models, SemStruct keeps the language model frozen and trains only a lightweight structural encoder. Extensive experiments on the Valentine and SOTAB-SM benchmarks demonstrate that SemStruct achieves state-of-the-art performance, outperforming fully fine-tuned baselines on complex, semantically joinable datasets. Furthermore, our ablation studies reveal that row representations serve primarily as topological conduits rather than semantic entities, validating the necessity of explicit structural modeling in schema matching.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on12
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan et al.VLDB 2021 · 484 citations
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- TAPEX: Table Pre-training via Learning a Neural SQL ExecutorQian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi et al.ICLR 2022 · 347 citations
Related papers
- TabEmb: Joint Semantic-Structure Embedding for Table AnnotationEhsan Hoseinzade, Ke Wang, Anandharaju Durai RajuACL 2026
- HeGTa: Leveraging Heterogeneous Graph-enhanced Large Language Models for Few-shot Complex Table UnderstandingRihui Jin, Yu Li, Guilin Qi, Nan Hu et al.AAAI 2025 · 1 citation
- In Situ Neural Relational Schema MatcherXingyu Du, Gongsheng Yuan, Sai Wu, Gang Chen et al.ICDE 2024 · 4 citations
- Translating Headers of Tabular Data: A Pilot Study of Schema TranslationKunrui Zhu, Yan Gao, Jiaqi Guo, Jian-Guang LouEMNLP 2021 · 2 citations
- Learning Contextual Representations for Semantic Parsing with Generation-Augmented Pre-TrainingPeng Shi, Patrick Ng, Zhiguo Wang, Henghui Zhu et al.AAAI 2021 · 124 citations
