ZTab: Domain-Based Zero-Shot Annotation for Table Columns
Ehsan Hoseinzade, Ke Wang
摘要
This study addresses the challenge of automatically detecting semantic column types in relational tables, a key task in many real-world applications. Zero-shot modeling eliminates the need for user-provided labeled training data, making it ideal for scenarios where data collection is costly or restricted due to issues such as privacy concerns. However, existing zeroshot models suffer from poor performance in the case of a large number of semantic column types or classes, poor understanding of tabular structures, and privacy risks arising from dependency on high-performance closed-source LLMs. We introduce ZTab, a domain-based zero-shot framework, to address both performance and zero-shot requirements. ZTab considers a domain configuration given by a set of predefined semantic types, plus sample table schemas based on such types, fine-tunes an annotation LLM using pseudo-tables generated for sample table schemas. ZTab is domain-based zero-shot in that it does not depend on user-specific labeled training data; therefore, no retraining is needed for a test table coming from a similar domain. We describe three cases for domain-based zero-shot. The domain configuration of ZTab provides a trade-off between the extent of zero-shot and the annotation performance: for a "universal domain" that contains all semantic types, domainbased zero-shot will approach "pure" zero-shot; on the other hand, a "specialized domain" that contains semantic types for a specific application will enable better zero-shot performance within that domain. The source code and datasets are available at https://github.com/hoseinzadeehsan/ZTab.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu 等VLDB 2021 · 被引用 2,406 次
- Synthetic Data Generation with Large Language Models for Text Classification: Potential and LimitationsZhuoyan Li, Hangxiao Zhu, Zhuoran Lu, Ming YinEMNLP 2023 · 被引用 102 次
- ZeroGen: Efficient Zero-shot Learning via Dataset GenerationJiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu 等EMNLP 2022 · 被引用 96 次
相关 Paper
- ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language ModelsBenjamin Feuer, Yurong Liu, Chinmay Hegde, Juliana FreireVLDB 2024 · 被引用 33 次
- Large Scale Transfer Learning for Tabular Data via Language ModelingJosh Gardner, Juan C. Perdomo, Ludwig SchmidtNeurIPS 2024 · 被引用 103 次
- Label Annotation for Tabular Anomaly Detection with Large Language ModelsHaihong Zhao, Aochuan Chen, Miao Peng, Xiaolong Fan 等KDD 2026
- Leveraging Table Content for Zero-shot Text-to-SQL with Meta-LearningYongrui Chen, Xinnan Guo, Chaojie Wang, Jian Qiu 等AAAI 2021 · 被引用 11 次
- TAROT: Task-Adaptive Refinement of LLM-prior Graphs for Few-shot Tabular LearningRuxue Shi, Yili Wang, Mengnan Du, Hangting Ye 等KDD 2026
