ZTab: Domain-Based Zero-Shot Annotation for Table Columns
Ehsan Hoseinzade, Ke Wang
Abstract
This study addresses the challenge of automatically detecting semantic column types in relational tables, a key task in many real-world applications. Zero-shot modeling eliminates the need for user-provided labeled training data, making it ideal for scenarios where data collection is costly or restricted due to issues such as privacy concerns. However, existing zeroshot models suffer from poor performance in the case of a large number of semantic column types or classes, poor understanding of tabular structures, and privacy risks arising from dependency on high-performance closed-source LLMs. We introduce ZTab, a domain-based zero-shot framework, to address both performance and zero-shot requirements. ZTab considers a domain configuration given by a set of predefined semantic types, plus sample table schemas based on such types, fine-tunes an annotation LLM using pseudo-tables generated for sample table schemas. ZTab is domain-based zero-shot in that it does not depend on user-specific labeled training data; therefore, no retraining is needed for a test table coming from a similar domain. We describe three cases for domain-based zero-shot. The domain configuration of ZTab provides a trade-off between the extent of zero-shot and the annotation performance: for a "universal domain" that contains all semantic types, domainbased zero-shot will approach "pure" zero-shot; on the other hand, a "specialized domain" that contains semantic types for a specific application will enable better zero-shot performance within that domain. The source code and datasets are available at https://github.com/hoseinzadeehsan/ZTab.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bceab022-5fe5-4e97-b501-2641b41462ceCited by top-tier papers1
Ask how each one uses itBuilds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu et al.VLDB 2021 · 2,406 citations
- Synthetic Data Generation with Large Language Models for Text Classification: Potential and LimitationsZhuoyan Li, Hangxiao Zhu, Zhuoran Lu, Ming YinEMNLP 2023 · 102 citations
- ZeroGen: Efficient Zero-shot Learning via Dataset GenerationJiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu et al.EMNLP 2022 · 96 citations
Related papers
- ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language ModelsBenjamin Feuer, Yurong Liu, Chinmay Hegde, Juliana FreireVLDB 2024 · 33 citations
- Large Scale Transfer Learning for Tabular Data via Language ModelingJosh Gardner, Juan C. Perdomo, Ludwig SchmidtNeurIPS 2024 · 103 citations
- Label Annotation for Tabular Anomaly Detection with Large Language ModelsHaihong Zhao, Aochuan Chen, Miao Peng, Xiaolong Fan et al.KDD 2026
- Leveraging Table Content for Zero-shot Text-to-SQL with Meta-LearningYongrui Chen, Xinnan Guo, Chaojie Wang, Jian Qiu et al.AAAI 2021 · 11 citations
- TAROT: Task-Adaptive Refinement of LLM-prior Graphs for Few-shot Tabular LearningRuxue Shi, Yili Wang, Mengnan Du, Hangting Ye et al.KDD 2026
