Pytheas: Pattern-based Table Discovery in CSV Files
Christina Christodoulakis, Eric B. Munson, Moshe Gabel, Angela Demke Brown, Renée J. Miller
Abstract
CSV is a popular Open Data format widely used in a variety of domains for its simplicity and effectiveness in storing and disseminating data. Unfortunately, data published in this format often does not conform to strict specifications, making automated data extraction from CSV files a painful task. While table discovery from HTML pages or spreadsheets has been studied extensively, extracting tables from CSV files still poses a considerable challenge due to their loosely defined format and limited embedded metadata. In this work we lay out the challenges of discovering tables in CSV files, and propose Pytheas: a principled method for automatically classifying lines in a CSV file and discovering tables within it based on the intuition that tables maintain a coherency of values in each column. We evaluate our methods over two manually annotated data sets: 2000 CSV files sampled from four Canadian Open Data portals, and 2500 additional files sampled from Canadian, US, UK and Australian portals. Our comparison to state-of-the-art approaches shows that Pytheas is able to successfully discover tables with precision and recall of over 95.9% and 95.7% respectively, while current approaches achieve around 89.6% precision and 81.3% recall. Furthermore, Pytheas's accuracy for correctly classifying all lines per CSV file is 95.6%, versus a maximum of 86.9% for compared approaches. Pytheas generalizes well to new data, with a table discovery Fmeasure above 95% even when trained on Canadian data and applied to data from different countries. Finally, we introduce a confidence measure for table discovery and demonstrate its value for accurately identifying potential errors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 54017fe2-c8a3-4214-bde4-5e125d6db475Cited by top-tier papers5
- Olio: A Semantic Search Interface for Data RepositoriesVidya Setlur, Andriy Kanyuka, Arjun SrinivasanUIST 2023 · 17 citations
- Pollock: A Data Loading BenchmarkGerardo Vitagliano, Mazhar Hameed, Lan Jiang, Lucas Reisener et al.VLDB 2023 · 9 citations
- Detecting Layout Templates in Complex Multiregion FilesGerardo Vitagliano, Lan Jiang, Felix NaumannVLDB 2022 · 7 citations
- Encoding Spreadsheets for Large Language ModelsHaoyu Dong, Jianbo Zhao, Yuzhang Tian, Junyu Xiong et al.EMNLP 2024 · 4 citations
- StructVizor: Interactive Profiling of Semi-Structured Textual DataYanwei Huang, Yan Miao, Di Weng, Adam Perer et al.CHI 2025 · 3 citations
Related papers
- Auto-Tables: Synthesizing Multi-Step Transformations to Relationalize Tables without Using ExamplesPeng Li, Yeye He, Cong Yan, Yue Wang et al.VLDB 2023 · 29 citations
- Relational Header Discovery using Similarity Search in a Table CorpusHazar Harmouch, Thorsten Papenbrock, Felix NaumannICDE 2021 · 6 citations
- AutoDDG: Automated Dataset Description Generation using Large Language ModelsHaoxiang Zhang, Yurong Liu, Aécio S. R. Santos, Wei-Lun Hung et al.SIGMOD 2026 · 17 citations
- Structure interpretation of text formatsSumit Gulwani, Vu Le, Arjun Radhakrishna, Ivan Radicek et al.OOPSLA 2020 · 1 citation
- Efficient Joinable Table Discovery in Data Lakes: A High-Dimensional Similarity-Based ApproachYuyang Dong, Kunihiro Takeoka, Chuan Xiao, Masafumi OyamadaICDE 2021 · 78 citations
