An Empirical Evaluation of Columnar Storage Formats
Xinyu Zeng, Yulong Hui, Jiahong Shen, Andrew Pavlo, Wes McKinney, Huanchen Zhang
Abstract
Columnar storage is a core component of a modern data analytics system. Although many database management systems (DBMSs) have proprietary storage formats, most provide extensive support to open-source storage formats such as Parquet and ORC to facilitate cross-platform data sharing. But these formats were developed over a decade ago, in the early 2010s, for the Hadoop ecosystem. Since then, both the hardware and workload landscapes have changed. In this paper, we revisit the most widely adopted open-source columnar storage formats (Parquet and ORC) with a deep dive into their internals. We designed a benchmark to stress-test the formats' performance and space efficiency under different workload configurations. From our comprehensive evaluation of Parquet and ORC, we identify design decisions advantageous with modern hardware and real-world data distributions. These include using dictionary encoding by default, favoring decoding speed over compression ratio for integer encoding algorithms, making block compression optional, and embedding finer-grained auxiliary data structures. We also point out the inefficiencies in the format designs when handling common machine learning workloads and using GPUs for decoding. Our analysis identified important considerations that may guide future formats to better fit modern technology trends.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 039ec2d0-8f1a-42c8-83b7-8fbaadac92a2Cited by top-tier papers17
- A Deep Dive into Common Open Formats for Analytical DBMSsChunwei Liu, Anna Pavlenko, Matteo Interlandi, Brandon HaynesVLDB 2023 · 23 citations
- Two Birds With One Stone: Designing a Hybrid Cloud Storage Engine for HTAPTobias Schmidt, Dominik Durner, Viktor Leis, Thomas NeumannVLDB 2024 · 12 citations
- F3: The Open-Source Data File Format for the FutureXinyu Zeng, Ruijun Meng, Martin Prammer, Wes McKinney et al.SIGMOD 2026 · 10 citations
- The FastLanes File FormatAzim Afroozeh, Peter BonczVLDB 2025 · 9 citations
- Femur: A Flexible Framework for Fast and Secure Querying from Public Key-Value StoreJiaoyi Zhang, Liqiang Peng, Mo Sha, Weiran Liu et al.SIGMOD 2025 · 4 citations
Builds on14
- A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database AnalyticsAnil Shanbhag, Samuel Madden, Xiangyao YuSIGMOD 2020 · 112 citations
- Chimp: Efficient Lossless Floating Point Compression for Time Series DatabasesPanagiotis Liakos, Katia Papakonstantinopoulou, Yannis KotidisVLDB 2022 · 76 citations
- DSB: A Decision Support Benchmark for Workload-Driven and Traditional Database SystemsBailu Ding, Surajit Chaudhuri, Johannes Gehrke, Vivek R. NarasayyaVLDB 2021 · 62 citations
- BtrBlocks: Efficient Columnar Compression for Data LakesMaximilian Kuschewski, David Sauerwein, Adnan Alhomssi, Viktor LeisSIGMOD 2023 · 47 citations
- Tile-based Lightweight Integer Compression in GPUAnil Shanbhag, Bobbi W. Yogatama, Xiangyao Yu, Samuel MaddenSIGMOD 2022 · 45 citations
Related papers
- Good to the Last Bit: Data-Driven Encoding with CodecDBHao Jiang, Chunwei Liu, John Paparrizos, Andrew A. Chien et al.SIGMOD 2021 · 45 citations
- Mainlining Databases: Supporting Fast Transactional Workloads on Universal Columnar Data File FormatsTianyu Li, Matthew Butrovich, Amadou Ngom, Wan Shen Lim et al.VLDB 2021 · 28 citations
- Nested Parquet Is Flat, Why Not Use It? How To Scan Nested Data With On-the-Fly Key Generation and JoinsAlice Rey, Maximilian Rieger, Thomas NeumannSIGMOD 2025 · 1 citation
- LST-Bench: Benchmarking Log-Structured Tables in the CloudJesús Camacho-Rodríguez, Ashvin Agrawal, Anja Gruenheid, Ashit Gosalia et al.SIGMOD 2024 · 7 citations
- The FastLanes Compression Layout: Decoding >100 Billion Integers per Second with Scalar CodeAzim Afroozeh, Peter BonczVLDB 2023 · 44 citations
