RECA: Related Tables Enhanced Column Semantic Type Annotation Framework
Yushi Sun, Hao Xin, Lei Chen
Abstract
Understanding the semantics of tabular data is of great importance in various downstream applications, such as schema matching, data cleaning, and data integration. Column semantic type annotation is a critical task in the semantic understanding of tabular data. Despite the fact that various approaches have been proposed, they are challenged by the difficulties of handling wide tables and incorporating complex inter-table context information. Failure to handle wide tables limits the usage of column type annotation approaches, while failure to incorporate inter-table context harms the annotation quality. Existing methods either completely ignore these problems or propose ad-hoc solutions. In this paper, we propose Related tables Enhanced Column semantic type Annotation framework (RECA), which incorporates inter-table context information by finding and aligning schema-similar and topic-relevant tables based on a novel named entity schema. The design of RECA can naturally handle wide tables and incorporate useful inter-table context information to enhance the annotation quality. We conduct extensive experiments on two web table datasets to comprehensively evaluate the performance of RECA. Our results show that RECA achieves support-weighted F1 scores of 0.853 and 0.937 with macro average F1 scores of 0.674 and 0.783 on the two datasets respectively, which outperform the state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 865130b6-b46c-49e5-96bf-7ca2d941e4fcCited by top-tier papers10
- Are Large Language Models a Good Replacement of Taxonomies?Yushi Sun, Xin Hao, Kai Sun, Yifan Xu et al.VLDB 2024 · 26 citations
- APEX-SQL: Talking to the data via Agentic Exploration for Text-to-SQLBowen Cao, Weibin Liao, Yushi Sun, Dong Fang et al.KDD 2026 · 7 citations
- KGLink: A Column Type Annotation Method that Combines Knowledge Graph and Pre-Trained Language ModelYubo Wang, Hao Xin, Lei ChenICDE 2024 · 7 citations
- Qualitative Join Discovery in Data Lakes using ExamplesMir Mahathir Mohammad, El Kindi RezigSIGMOD 2026 · 6 citations
- Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in TablesQixu Chen, Yeye He, Raymond Chi-Wing Wong, Weiwei Cui et al.SIGMOD 2025 · 4 citations
Builds on6
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu et al.VLDB 2021 · 2,406 citations
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- Annotating Columns with Pre-trained Language ModelsYoshihiko Suhara, Jinfeng Li, Yuliang Li, Dan Zhang et al.SIGMOD 2022 · 81 citations
- TCN: Table Convolutional Network for Web Table InterpretationDaheng Wang, Prashant Shiralkar, Colin Lockard, Binxuan Huang et al.WWW 2021 · 68 citations
- TaPas: Weakly Supervised Table Parsing via Pre-trainingJonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller, Francesco Piccinno et al.ACL 2020 · 19 citations
Related papers
- Retrieve-and-Verify: A Table Context Selection Framework for Accurate Column AnnotationsZhihao Ding, Yongkang Sun, Jieming ShiSIGMOD 2026 · 2 citations
- Sato: Contextual Semantic Type Detection in TablesDan Zhang, Yoshihiko Suhara, Jinfeng Li, Madelon Hulsebos et al.VLDB 2020
- Watchog: A Light-weight Contrastive Learning based Framework for Column AnnotationZhengjie Miao, Jin WangSIGMOD 2024 · 14 citations
- Relational Header Discovery using Similarity Search in a Table CorpusHazar Harmouch, Thorsten Papenbrock, Felix NaumannICDE 2021 · 6 citations
- CARTE: Pretraining and Transfer for Tabular LearningMyung Jun Kim, Léo Grinsztajn, Gaël VaroquauxICML 2024 · 52 citations
