Annotating Columns with Pre-trained Language Models
Yoshihiko Suhara, Jinfeng Li, Yuliang Li, Dan Zhang, Çagatay Demiralp, Chen Chen, Wang-Chiew Tan
摘要
Inferring meta information about tables, such as column headers or relationships between columns, is an active research topic in data management as we find many tables are missing some of this information. In this paper, we study the problem of annotating table columns (i.e., predicting column types and the relationships between columns) using only information from the table itself. We develop a multi-task learning framework (called Doduo) based on pre-trained language models, which takes the entire table as input and predicts column types/relations using a single model. Experimental results show that Doduo establishes new state-of-the-art performance on two benchmarks for the column type prediction and column relation prediction tasks with up to 4.0% and 11.9% improvements, respectively. We report that Doduo can already outperform the previous state-of-the-art performance with a minimal number of tokens, only 8 tokens per column. We release a toolbox (https://github.com/megagonlabs/doduo) and confirm the effectiveness of Doduo on a real-world data science problem through a case study.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation LearningGrace Fan, Jin Wang, Yuliang Li, Dan Zhang 等VLDB 2023 · 被引用 139 次
- HyTrel: Hypergraph-enhanced Tabular Data Representation LearningPei Chen, Soumajyoti Sarkar, Leonard Lausen, Balasubramaniam Srinivasan 等NeurIPS 2023 · 被引用 66 次
- Table-GPT: Table Fine-tuned GPT for Diverse Table TasksPeng Li, Yeye He, Dror Yashar, Weiwei Cui 等SIGMOD 2024 · 被引用 63 次
- SANTOS: Relationship-based Semantic Table Union SearchAamod Khatiwada, Grace Fan, Roee Shraga, Zixuan Chen 等SIGMOD 2023 · 被引用 61 次
- Integrating Data Lake TablesAamod Khatiwada, Roee Shraga, Wolfgang Gatterbauer, Renée J. MillerVLDB 2023 · 被引用 59 次
它引用的顶会 Paper12
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu 等VLDB 2021 · 被引用 2,406 次
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan 等VLDB 2021 · 被引用 484 次
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 被引用 417 次
- Creating Embeddings of Heterogeneous Relational Datasets for Data Integration TasksRiccardo Cappuzzo, Paolo Papotti, Saravanan ThirumuruganathanSIGMOD 2020 · 被引用 139 次
- RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data PreparationNan Tang, Ju Fan, Fangyi Li, Jianhong Tu 等VLDB 2021 · 被引用 92 次
相关 Paper
- Label-Constrained Column Annotation with Language Models and Graph Neural NetworksDuo Yang, Ioannis Dasoulas, Anastasia DimouICDE 2026
- Retrieve-and-Verify: A Table Context Selection Framework for Accurate Column AnnotationsZhihao Ding, Yongkang Sun, Jieming ShiSIGMOD 2026 · 被引用 2 次
- Watchog: A Light-weight Contrastive Learning based Framework for Column AnnotationZhengjie Miao, Jin WangSIGMOD 2024 · 被引用 14 次
- ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language ModelsBenjamin Feuer, Yurong Liu, Chinmay Hegde, Juliana FreireVLDB 2024 · 被引用 33 次
- TabEmb: Joint Semantic-Structure Embedding for Table AnnotationEhsan Hoseinzade, Ke Wang, Anandharaju Durai RajuACL 2026
