GIO: Generating Efficient Matrix and Frame Readers for Custom Data Formats by Example
Saeed Fathollahzadeh, Matthias Boehm
Abstract
Data Scientists deal with a wide variety of file data formats and data representations. Probably the most difficult to handle are custom data formats that liberally define their own particular flat or nested structure with multiple custom delimiters, multi-line records, or undocumented semantics of attribute sequences, coappearances, and repetitions. As a prerequisite for exploratory ML model training, data scientists need to map these data representations into regular frames or matrices. Unfortunately, existing tools and frameworks provide only limited support for aiding this process, which causes redundant manual efforts and unnecessary data quality issues. In this paper, we initiate work on automatic matrix and frame reader generation by example. A user provides a sample of raw text data and its mapped matrix or frame representation. Our GIO framework then first identifies the mapping rules from raw to structured data, and subsequently generates source code of an efficient, multi-threaded reader for reading full raw datasets of this format. In order to facilitate manual improvements, both the mapping rules, and generated reader can be modified as needed. Our experiments show that GIO is able to correctly identify the mapping rules for basic text formats like CSV, LibSVM, MatrixMarket; custom text formats from publishing, automotive, and health care; as well as various nested formats such as JSON and XML. Additionally, the automatically generated readers yield competitive performance compared to hand-coded readers and tuned libraries like RapidJSON. CCS Concepts: • Information systems → Database management system engines; Data scans; Record and block layout; • Theory of computation → Database query processing and optimization (theory).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f29de7e-b016-4bf8-b5b2-f488ce40aaf9Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Grizzly: Efficient Stream Processing Through Adaptive Query CompilationPhilipp M. Grulich, Sebastian Breß, Steffen Zeuch, Jonas Traub et al.SIGMOD 2020 · 41 citations
- JSON Tiles: Fast Analytics on Semi-Structured DataDominik Durner, Viktor Leis, Thomas NeumannSIGMOD 2021 · 28 citations
- Towards Benchmarking Feature Type Inference for AutoML PlatformsVraj Shah, Jonathan Lacanlale, Premanand Kumar, Kevin Yang et al.SIGMOD 2021 · 16 citations
- Scalable Structural Index Construction for JSON AnalyticsLin Jiang, Junqiao Qiu, Zhijia ZhaoVLDB 2021 · 16 citations
- Witness Generation for JSON SchemaLyes Attouche, Mohamed-Amine Baazizi, Dario Colazzo, Giorgio Ghelli et al.VLDB 2022 · 14 citations
Related papers
- MojoFrame: Dataframe Library in Mojo LanguageShengya Huang, Zhaoheng Li, Derek Werner, Yongjoo ParkICDE 2026
- LiteForm: Lightweight and Automatic Format Composition for Sparse Matrix-Matrix Multiplication on GPUsZhen Peng, Polykarpos Thomadakis, Jacques A. Pienaar, Gokcen KestorHPDC 2025 · 1 citation
- DTT: An Example-Driven Tabular Transformer for Joinability by Leveraging Large Language ModelsArash Dargahi Nobari, Davood RafieiSIGMOD 2024 · 11 citations
- Pollock: A Data Loading BenchmarkGerardo Vitagliano, Mazhar Hameed, Lan Jiang, Lucas Reisener et al.VLDB 2023 · 9 citations
- Automatic Semantic Alignment of Flow Pattern Representations for Exploration with Large Language ModelsWeihan Zhang, Jun TaoIEEE VIS 2025 · 1 citation
