Pytheas: Pattern-based Table Discovery in CSV Files
Christina Christodoulakis, Eric B. Munson, Moshe Gabel, Angela Demke Brown, Renée J. Miller
摘要
CSV is a popular Open Data format widely used in a variety of domains for its simplicity and effectiveness in storing and disseminating data. Unfortunately, data published in this format often does not conform to strict specifications, making automated data extraction from CSV files a painful task. While table discovery from HTML pages or spreadsheets has been studied extensively, extracting tables from CSV files still poses a considerable challenge due to their loosely defined format and limited embedded metadata. In this work we lay out the challenges of discovering tables in CSV files, and propose Pytheas: a principled method for automatically classifying lines in a CSV file and discovering tables within it based on the intuition that tables maintain a coherency of values in each column. We evaluate our methods over two manually annotated data sets: 2000 CSV files sampled from four Canadian Open Data portals, and 2500 additional files sampled from Canadian, US, UK and Australian portals. Our comparison to state-of-the-art approaches shows that Pytheas is able to successfully discover tables with precision and recall of over 95.9% and 95.7% respectively, while current approaches achieve around 89.6% precision and 81.3% recall. Furthermore, Pytheas's accuracy for correctly classifying all lines per CSV file is 95.6%, versus a maximum of 86.9% for compared approaches. Pytheas generalizes well to new data, with a table discovery Fmeasure above 95% even when trained on Canadian data and applied to data from different countries. Finally, we introduce a confidence measure for table discovery and demonstrate its value for accurately identifying potential errors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Olio: A Semantic Search Interface for Data RepositoriesVidya Setlur, Andriy Kanyuka, Arjun SrinivasanUIST 2023 · 被引用 17 次
- Pollock: A Data Loading BenchmarkGerardo Vitagliano, Mazhar Hameed, Lan Jiang, Lucas Reisener 等VLDB 2023 · 被引用 9 次
- Detecting Layout Templates in Complex Multiregion FilesGerardo Vitagliano, Lan Jiang, Felix NaumannVLDB 2022 · 被引用 7 次
- Encoding Spreadsheets for Large Language ModelsHaoyu Dong, Jianbo Zhao, Yuzhang Tian, Junyu Xiong 等EMNLP 2024 · 被引用 4 次
- StructVizor: Interactive Profiling of Semi-Structured Textual DataYanwei Huang, Yan Miao, Di Weng, Adam Perer 等CHI 2025 · 被引用 3 次
相关 Paper
- Auto-Tables: Synthesizing Multi-Step Transformations to Relationalize Tables without Using ExamplesPeng Li, Yeye He, Cong Yan, Yue Wang 等VLDB 2023 · 被引用 29 次
- Relational Header Discovery using Similarity Search in a Table CorpusHazar Harmouch, Thorsten Papenbrock, Felix NaumannICDE 2021 · 被引用 6 次
- AutoDDG: Automated Dataset Description Generation using Large Language ModelsHaoxiang Zhang, Yurong Liu, Aécio S. R. Santos, Wei-Lun Hung 等SIGMOD 2026 · 被引用 17 次
- Structure interpretation of text formatsSumit Gulwani, Vu Le, Arjun Radhakrishna, Ivan Radicek 等OOPSLA 2020 · 被引用 1 次
- Efficient Joinable Table Discovery in Data Lakes: A High-Dimensional Similarity-Based ApproachYuyang Dong, Kunihiro Takeoka, Chuan Xiao, Masafumi OyamadaICDE 2021 · 被引用 78 次
