GitTables: A Large-Scale Corpus of Relational Tables
Madelon Hulsebos, Çagatay Demiralp, Paul Groth
Abstract
The success of deep learning has sparked interest in improving relational table tasks, like data preparation and search, with table representation models trained on large table corpora. Existing table corpora primarily contain tables extracted from HTML pages, limiting the capability to represent offline database tables. To train and evaluate high-capacity models for applications beyond the Web, we need resources with tables that resemble relational database tables. Here we introduce GitTables, a corpus of 1M relational tables extracted from GitHub. Our continuing curation aims at growing the corpus to at least 10M tables. Analyses of GitTables show that its structure, content, and topical coverage differ significantly from existing table corpora. We annotate table columns in GitTables with semantic types, hierarchical relations and descriptions from Schema.org and DBpedia. The evaluation of our annotation pipeline on the T2Dv2 benchmark illustrates that our approach provides results on par with human annotations. We present three applications of GitTables, demonstrating its value for learned semantic type detection models, schema completion methods, and benchmarks for table-to-KG matching, data search, and preparation. We make the corpus and code available at https://gittables.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 90843090-cc7e-4990-b1cf-2fbe3d8bcd05Cited by top-tier papers26
- OmniSQL: Synthesizing High-quality Text-to-SQL Data at ScaleHaoyang Li, Shang Wu, Xiaokang Zhang, Xinmei Huang et al.VLDB 2025 · 90 citations
- CARTE: Pretraining and Transfer for Tabular LearningMyung Jun Kim, Léo Grinsztajn, Gaël VaroquauxICML 2024 · 52 citations
- ArcheType: A Novel Framework for Open-Source Column Type Annotation using Large Language ModelsBenjamin Feuer, Yurong Liu, Chinmay Hegde, Juliana FreireVLDB 2024 · 33 citations
- CHORUS: Foundation Models for Unified Data Discovery and ExplorationMoe Kayali, Anton Lykov, Ilias Fountalis, Nikolaos Vasiloglou et al.VLDB 2024 · 33 citations
- Towards Cross-Table Masked Pretraining for Web Data MiningChao Ye, Guoshan Lu, Haobo Wang, Liyao Li et al.WWW 2024 · 23 citations
Builds on4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu et al.VLDB 2021 · 2,406 citations
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 417 citations
- TCN: Table Convolutional Network for Web Table InterpretationDaheng Wang, Prashant Shiralkar, Colin Lockard, Binxuan Huang et al.WWW 2021 · 68 citations
Related papers
- SchemaPile: A Large Collection of Relational Database SchemasTill Döhmen, Radu Geacu, Madelon Hulsebos, Sebastian SchelterSIGMOD 2024 · 9 citations
- Watchog: A Light-weight Contrastive Learning based Framework for Column AnnotationZhengjie Miao, Jin WangSIGMOD 2024 · 14 citations
- KTabulator: Interactive Ad hoc Table Creation using Knowledge GraphsSiyuan Xia, Nafisa Anzum, Semih Salihoglu, Jian ZhaoCHI 2021 · 7 citations
- Sato: Contextual Semantic Type Detection in TablesDan Zhang, Yoshihiko Suhara, Jinfeng Li, Madelon Hulsebos et al.VLDB 2020
- TempTabQA: Temporal Question Answering for Semi-Structured TablesVivek Gupta, Pranshu Kandoi, Mahek Bhavesh Vora, Shuo Zhang et al.EMNLP 2023 · 4 citations
