SchemaPile: A Large Collection of Relational Database Schemas
Till Döhmen, Radu Geacu, Madelon Hulsebos, Sebastian Schelter
摘要
Access to fine-grained schema information is crucial for understanding how relational databases are designed and used in practice, and for building systems that help users interact with them. Furthermore, such information is required as training data to leverage the potential of large language models (LLMs) for improving data preparation, data integration and natural language querying. Existing single-table corpora such as GitTables provide insights into how tables are structured in-the-wild, but lack detailed schema information about how tables relate to each other, as well as metadata like data types or integrity constraints. On the other hand, existing multi-table (or database schema) datasets are rather small and attribute-poor, leaving it unclear to what extent they actually represent typical real-world database schemas.
In order to address these challenges, we present SchemaPile, a corpus of 221,171 database schemas, extracted from SQL files on GitHub. It contains 1.7 million tables with 10 million column definitions, 700 thousand foreign key relationships, seven million integrity constraints, and data content for more than 340 thousand tables. We conduct an in-depth analysis on the millions of schema metadata properties in our corpus, as well as its highly diverse language and topic distribution. In addition, we showcase the potential of SchemaPile to improve a variety of data management applications, e.g., fine-tuning LLMs for schema-only foreign key detection, improving CSV header detection and evaluating multi-dialect SQL parsers. We publish the code and data for recreating SchemaPile and a permissively licensed subset SchemaPile-Perm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- OmniSQL: Synthesizing High-quality Text-to-SQL Data at ScaleHaoyang Li, Shang Wu, Xiaokang Zhang, Xinmei Huang 等VLDB 2025 · 被引用 90 次
- SNAILS: Schema Naming Assessments for Improved LLM-Based SQL InferenceKyle Luoma, Arun KumarSIGMOD 2025 · 被引用 11 次
- Qualitative Join Discovery in Data Lakes using ExamplesMir Mahathir Mohammad, El Kindi RezigSIGMOD 2026 · 被引用 6 次
- WikiDBGraph: A Data Management Benchmark Suite for Collaborative Learning Over Database SilosZhaomin Wu, Ziyang Wang, Bingsheng HeICDE 2026
- Dialect-Agnostic SQL Parsing via LLM-Based SegmentationJunwen An, Kabilan Mahathevan, Manuel RiggerSIGMOD 2026
它引用的顶会 Paper11
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- How Large Language Models Will Disrupt Data ManagementRaul Castro Fernandez, Aaron J. Elmore, Michael J. Franklin, Sanjay Krishnan 等VLDB 2023 · 被引用 127 次
- Valentine: Evaluating Matching Techniques for Dataset DiscoveryChristos Koutras, George Siachamis, Andra Ionescu, Kyriakos Psarakis 等ICDE 2021 · 被引用 87 次
- Few-shot Text-to-SQL Translation using Structure and Content Prompt LearningZihui Gu, Ju Fan, Nan Tang, Lei Cao 等SIGMOD 2023 · 被引用 60 次
- GraPPa: Grammar-Augmented Pre-Training for Table Semantic ParsingTao Yu, Chien-Sheng Wu, Xi Victoria Lin, Bailin Wang 等ICLR 2021 · 被引用 59 次
相关 Paper
- GitTables: A Large-Scale Corpus of Relational TablesMadelon Hulsebos, Çagatay Demiralp, Paul GrothSIGMOD 2023 · 被引用 42 次
- Can Large Language Models Predict Data Correlations from Column Names?Immanuel TrummerVLDB 2023 · 被引用 17 次
- GRIT: Guided Relational Integration for Efficient Multi-Table UnderstandingYujin Kang, Park Seong Woo, Yoon-Sik ChoEMNLP 2025 · 被引用 2 次
- SchemaRAG: A Schema-aware Retrieval-Augmented Generation Framework for Text-to-SQLDi Wu, Zetong Tang, Yi He, Xin LuoSIGMOD 2026 · 被引用 9 次
- In Situ Neural Relational Schema MatcherXingyu Du, Gongsheng Yuan, Sai Wu, Gang Chen 等ICDE 2024 · 被引用 4 次
