Pollock: A Data Loading Benchmark
Gerardo Vitagliano, Mazhar Hameed, Lan Jiang, Lucas Reisener, Eugene Wu, Felix Naumann
Abstract
Any system at play in a data-driven project has a fundamental requirement: the ability to load data. The de-facto standard format to distribute and consume raw data is csv. Yet, the plain text and flexible nature of this format make such files often difficult to parse and correctly load their content, requiring cumbersome data preparation steps. We propose a benchmark to assess the robustness of systems in loading data from non-standard csv formats and with structural inconsistencies. First, we formalize a model to describe the issues that affect real-world files and use it to derive a systematic "pollution" process to generate dialects for any given grammar. Our benchmark leverages the pollution framework for the csv format. To guide pollution, we have surveyed thousands of real-world, publicly available csv files, recording the problems we encountered. We demonstrate the applicability of our benchmark by testing and scoring 16 different systems: popular csv parsing frameworks, relational database tools, spreadsheet systems, and a data visualization tool.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7b1a0edb-bec3-47a5-9307-b6136996c0bcCited by top-tier papers2
- Table-GPT: Table Fine-tuned GPT for Diverse Table TasksPeng Li, Yeye He, Dror Yashar, Weiwei Cui et al.SIGMOD 2024 · 63 citations
- SchemaPile: A Large Collection of Relational Database SchemasTill Döhmen, Radu Geacu, Madelon Hulsebos, Sebastian SchelterSIGMOD 2024 · 9 citations
Builds on4
- BigFuzz: Efficient Fuzz Testing for Data Analytics Using Framework AbstractionQian Zhang, Jiyuan Wang, Muhammad Ali Gulzar, Rohan Padhye et al.ASE 2020 · 27 citations
- Towards Benchmarking Feature Type Inference for AutoML PlatformsVraj Shah, Jonathan Lacanlale, Premanand Kumar, Kevin Yang et al.SIGMOD 2021 · 16 citations
- Detecting Layout Templates in Complex Multiregion FilesGerardo Vitagliano, Lan Jiang, Felix NaumannVLDB 2022 · 7 citations
- Pytheas: Pattern-based Table Discovery in CSV FilesChristina Christodoulakis, Eric B. Munson, Moshe Gabel, Angela Demke Brown et al.VLDB 2020
Related papers
- One Pass to Parse Them All: Fused Parallel CSV ProcessingSimon Ellmann, Thomas NeumannVLDB 2026
- Are we building on the rock? on the importance of data preprocessing for code summarizationLin Shi, Fangwen Mu, Xiao Chen, Song Wang et al.FSE 2022 · 65 citations
- Mind the Gap: An Experimental Evaluation of Imputation of Missing Values Techniques in Time SeriesMourad Khayati, Alberto Lerner, Zakhar Tymchenko, Philippe Cudré-MaurouxVLDB 2020 · 57 citations
- PU-BENCH: A Unified Benchmark for Rigorous and Reproducible PU LearningQiuyi Chen, Haiyang Zhang, Leqi Zhang, Changchun Li et al.ICLR 2026
- Auto-Tables: Synthesizing Multi-Step Transformations to Relationalize Tables without Using ExamplesPeng Li, Yeye He, Cong Yan, Yue Wang et al.VLDB 2023 · 29 citations
