KGLink: A Column Type Annotation Method that Combines Knowledge Graph and Pre-Trained Language Model
Yubo Wang, Hao Xin, Lei Chen
Abstract
The semantic annotation of tabular data plays a crucial role in various downstream tasks. Previous research has proposed knowledge graph (KG)-based and deep learning-based methods, each with its inherent limitations. KG-based methods encounter difficulties annotating columns when there is no match for column cells in the KG. Moreover, KG-based methods can provide multiple predictions for one column, making it challenging to determine the semantic type with the most suitable granularity for the dataset. This type granularity issue limits their scalability. On the other hand, deep learning-based methods face challenges related to the valuable context missing issue. This occurs when the information within the table is insufficient for determining the correct column type. This paper presents KGLink, a method that combines Wiki-Data KG information with a pre-trained deep learning language model for table column annotation, effectively addressing both type granularity and valuable context missing issues. Through comprehensive experiments on widely used tabular datasets encompassing numeric and string columns with varying type granularity, we showcase the effectiveness and efficiency of KGLink. By leveraging the strengths of KGLink, we successfully surmount challenges related to type granularity and valuable context issues, establishing it as a robust solution for the semantic annotation of tabular data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dce2c454-c1e4-4c75-8878-22592f400e25Cited by top-tier papers4
- Qualitative Join Discovery in Data Lakes using ExamplesMir Mahathir Mohammad, El Kindi RezigSIGMOD 2026 · 6 citations
- LLMLog: Advanced Log Template Generation via LLM-driven Multi-Round AnnotationFei Teng, Haoyang Li, Lei ChenVLDB 2025 · 2 citations
- ZTab: Domain-Based Zero-Shot Annotation for Table ColumnsEhsan Hoseinzade, Ke WangICDE 2026
- TabEmb: Joint Semantic-Structure Embedding for Table AnnotationEhsan Hoseinzade, Ke Wang, Anandharaju Durai RajuACL 2026
Builds on10
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu et al.VLDB 2021 · 2,406 citations
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- What Language Model Architecture and Pretraining Objective Works Best for Zero-Shot Generalization?Thomas Wang, Adam Roberts, Daniel Hesslow, Teven Le Scao et al.ICML 2022 · 228 citations
- Annotating Columns with Pre-trained Language ModelsYoshihiko Suhara, Jinfeng Li, Yuliang Li, Dan Zhang et al.SIGMOD 2022 · 81 citations
Related papers
- RECA: Related Tables Enhanced Column Semantic Type Annotation FrameworkYushi Sun, Hao Xin, Lei ChenVLDB 2023 · 26 citations
- Retrieve-and-Verify: A Table Context Selection Framework for Accurate Column AnnotationsZhihao Ding, Yongkang Sun, Jieming ShiSIGMOD 2026 · 2 citations
- MixedSAND: Semantic Annotation of Mixed-unit Numeric DataAmir Behrad Khorram Nazari, Davood Rafiei, Mario A. NascimentoWWW 2025 · 1 citation
- Label-Constrained Column Annotation with Language Models and Graph Neural NetworksDuo Yang, Ioannis Dasoulas, Anastasia DimouICDE 2026
- Sato: Contextual Semantic Type Detection in TablesDan Zhang, Yoshihiko Suhara, Jinfeng Li, Madelon Hulsebos et al.VLDB 2020
