GIO: Generating Efficient Matrix and Frame Readers for Custom Data Formats by Example
Saeed Fathollahzadeh, Matthias Boehm
摘要
Data Scientists deal with a wide variety of file data formats and data representations. Probably the most difficult to handle are custom data formats that liberally define their own particular flat or nested structure with multiple custom delimiters, multi-line records, or undocumented semantics of attribute sequences, coappearances, and repetitions. As a prerequisite for exploratory ML model training, data scientists need to map these data representations into regular frames or matrices. Unfortunately, existing tools and frameworks provide only limited support for aiding this process, which causes redundant manual efforts and unnecessary data quality issues. In this paper, we initiate work on automatic matrix and frame reader generation by example. A user provides a sample of raw text data and its mapped matrix or frame representation. Our GIO framework then first identifies the mapping rules from raw to structured data, and subsequently generates source code of an efficient, multi-threaded reader for reading full raw datasets of this format. In order to facilitate manual improvements, both the mapping rules, and generated reader can be modified as needed. Our experiments show that GIO is able to correctly identify the mapping rules for basic text formats like CSV, LibSVM, MatrixMarket; custom text formats from publishing, automotive, and health care; as well as various nested formats such as JSON and XML. Additionally, the automatically generated readers yield competitive performance compared to hand-coded readers and tuned libraries like RapidJSON. CCS Concepts: • Information systems → Database management system engines; Data scans; Record and block layout; • Theory of computation → Database query processing and optimization (theory).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- Grizzly: Efficient Stream Processing Through Adaptive Query CompilationPhilipp M. Grulich, Sebastian Breß, Steffen Zeuch, Jonas Traub 等SIGMOD 2020 · 被引用 41 次
- JSON Tiles: Fast Analytics on Semi-Structured DataDominik Durner, Viktor Leis, Thomas NeumannSIGMOD 2021 · 被引用 28 次
- Towards Benchmarking Feature Type Inference for AutoML PlatformsVraj Shah, Jonathan Lacanlale, Premanand Kumar, Kevin Yang 等SIGMOD 2021 · 被引用 16 次
- Scalable Structural Index Construction for JSON AnalyticsLin Jiang, Junqiao Qiu, Zhijia ZhaoVLDB 2021 · 被引用 16 次
- Witness Generation for JSON SchemaLyes Attouche, Mohamed-Amine Baazizi, Dario Colazzo, Giorgio Ghelli 等VLDB 2022 · 被引用 14 次
相关 Paper
- MojoFrame: Dataframe Library in Mojo LanguageShengya Huang, Zhaoheng Li, Derek Werner, Yongjoo ParkICDE 2026
- LiteForm: Lightweight and Automatic Format Composition for Sparse Matrix-Matrix Multiplication on GPUsZhen Peng, Polykarpos Thomadakis, Jacques A. Pienaar, Gokcen KestorHPDC 2025 · 被引用 1 次
- DTT: An Example-Driven Tabular Transformer for Joinability by Leveraging Large Language ModelsArash Dargahi Nobari, Davood RafieiSIGMOD 2024 · 被引用 11 次
- Pollock: A Data Loading BenchmarkGerardo Vitagliano, Mazhar Hameed, Lan Jiang, Lucas Reisener 等VLDB 2023 · 被引用 9 次
- Automatic Semantic Alignment of Flow Pattern Representations for Exploration with Large Language ModelsWeihan Zhang, Jun TaoIEEE VIS 2025 · 被引用 1 次
