A Deep Dive into Common Open Formats for Analytical DBMSs
Chunwei Liu, Anna Pavlenko, Matteo Interlandi, Brandon Haynes
Abstract
This paper evaluates the suitability of Apache Arrow, Parquet, and ORC as formats for subsumption in an analytical DBMS. We systematically identify and explore the high-level features that are important to support efficient querying in modern OLAP DBMSs and evaluate the ability of each format to support these features. We find that each format has trade-offs that make it more or less suitable for use as a format in a DBMS and identify opportunities to more holistically co-design a unified in-memory and on-disk data representation. Our hope is that this study can be used as a guide for system developers designing and using these formats, as well as provide the community with directions to pursue for improving these common open formats.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8cb920b4-e6ab-4499-9ffc-d3bd636a37eaCited by top-tier papers10
- Two Birds With One Stone: Designing a Hybrid Cloud Storage Engine for HTAPTobias Schmidt, Dominik Durner, Viktor Leis, Thomas NeumannVLDB 2024 · 12 citations
- F3: The Open-Source Data File Format for the FutureXinyu Zeng, Ruijun Meng, Martin Prammer, Wes McKinney et al.SIGMOD 2026 · 10 citations
- The FastLanes File FormatAzim Afroozeh, Peter BonczVLDB 2025 · 9 citations
- Beyond Compression: A Comprehensive Evaluation of Lossless Floating-Point CompressionKaisei Hishida, Chunwei Liu, John Paparrizos, Aaron J. ElmoreVLDB 2025 · 8 citations
- AdaEdge: A Dynamic Compression Selection Framework for Resource Constrained DevicesChunwei Liu, John Paparrizos, Aaron J. ElmoreICDE 2024 · 7 citations
Builds on10
- Qd-tree: Learning Data Layouts for Big Data AnalyticsZongheng Yang, Badrish Chandramouli, Chi Wang, Johannes Gehrke et al.SIGMOD 2020 · 87 citations
- FlexPushdownDB: Hybrid Pushdown and Caching in a Cloud DBMSYifei Yang, Matt Youill, Matthew E. Woicik, Yizhou Liu et al.VLDB 2021 · 67 citations
- Decomposed Bounded Floats for Fast Compression and QueriesChunwei Liu, Hao Jiang, John Paparrizos, Aaron J. ElmoreVLDB 2021 · 65 citations
- An Empirical Evaluation of Columnar Storage FormatsXinyu Zeng, Yulong Hui, Jiahong Shen, Andrew Pavlo et al.VLDB 2024 · 59 citations
- CompressDB: Enabling Efficient Compressed Data Direct Processing for Various DatabasesFeng Zhang, Weitao Wan, Chenyang Zhang, Jidong Zhai et al.SIGMOD 2022 · 46 citations
Related papers
- Mainlining Databases: Supporting Fast Transactional Workloads on Universal Columnar Data File FormatsTianyu Li, Matthew Butrovich, Amadou Ngom, Wan Shen Lim et al.VLDB 2021 · 28 citations
- BOSS - An Architecture for Database Kernel CompositionHubert Mohr-Daurat, Xuan Sun, Holger PirkVLDB 2024 · 12 citations
- Nested Parquet Is Flat, Why Not Use It? How To Scan Nested Data With On-the-Fly Key Generation and JoinsAlice Rey, Maximilian Rieger, Thomas NeumannSIGMOD 2025 · 1 citation
- Maximizing Persistent Memory Bandwidth Utilization for OLAP WorkloadsBjörn Daase, Lars Jonas Bollmeier, Lawrence Benson, Tilmann RablSIGMOD 2021 · 38 citations
- Active Data Lakes: Regaining Physical Data Independence Without Losing InteroperabilityPascal Ginter, Viktor LeisVLDB 2026 · 4 citations
