An Empirical Evaluation of Columnar Storage Formats
Xinyu Zeng, Yulong Hui, Jiahong Shen, Andrew Pavlo, Wes McKinney, Huanchen Zhang
摘要
Columnar storage is a core component of a modern data analytics system. Although many database management systems (DBMSs) have proprietary storage formats, most provide extensive support to open-source storage formats such as Parquet and ORC to facilitate cross-platform data sharing. But these formats were developed over a decade ago, in the early 2010s, for the Hadoop ecosystem. Since then, both the hardware and workload landscapes have changed. In this paper, we revisit the most widely adopted open-source columnar storage formats (Parquet and ORC) with a deep dive into their internals. We designed a benchmark to stress-test the formats' performance and space efficiency under different workload configurations. From our comprehensive evaluation of Parquet and ORC, we identify design decisions advantageous with modern hardware and real-world data distributions. These include using dictionary encoding by default, favoring decoding speed over compression ratio for integer encoding algorithms, making block compression optional, and embedding finer-grained auxiliary data structures. We also point out the inefficiencies in the format designs when handling common machine learning workloads and using GPUs for decoding. Our analysis identified important considerations that may guide future formats to better fit modern technology trends.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- A Deep Dive into Common Open Formats for Analytical DBMSsChunwei Liu, Anna Pavlenko, Matteo Interlandi, Brandon HaynesVLDB 2023 · 被引用 23 次
- Two Birds With One Stone: Designing a Hybrid Cloud Storage Engine for HTAPTobias Schmidt, Dominik Durner, Viktor Leis, Thomas NeumannVLDB 2024 · 被引用 12 次
- F3: The Open-Source Data File Format for the FutureXinyu Zeng, Ruijun Meng, Martin Prammer, Wes McKinney 等SIGMOD 2026 · 被引用 10 次
- The FastLanes File FormatAzim Afroozeh, Peter BonczVLDB 2025 · 被引用 9 次
- Femur: A Flexible Framework for Fast and Secure Querying from Public Key-Value StoreJiaoyi Zhang, Liqiang Peng, Mo Sha, Weiran Liu 等SIGMOD 2025 · 被引用 4 次
它引用的顶会 Paper14
- A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database AnalyticsAnil Shanbhag, Samuel Madden, Xiangyao YuSIGMOD 2020 · 被引用 112 次
- Chimp: Efficient Lossless Floating Point Compression for Time Series DatabasesPanagiotis Liakos, Katia Papakonstantinopoulou, Yannis KotidisVLDB 2022 · 被引用 76 次
- DSB: A Decision Support Benchmark for Workload-Driven and Traditional Database SystemsBailu Ding, Surajit Chaudhuri, Johannes Gehrke, Vivek R. NarasayyaVLDB 2021 · 被引用 62 次
- BtrBlocks: Efficient Columnar Compression for Data LakesMaximilian Kuschewski, David Sauerwein, Adnan Alhomssi, Viktor LeisSIGMOD 2023 · 被引用 47 次
- Tile-based Lightweight Integer Compression in GPUAnil Shanbhag, Bobbi W. Yogatama, Xiangyao Yu, Samuel MaddenSIGMOD 2022 · 被引用 45 次
相关 Paper
- Good to the Last Bit: Data-Driven Encoding with CodecDBHao Jiang, Chunwei Liu, John Paparrizos, Andrew A. Chien 等SIGMOD 2021 · 被引用 45 次
- Mainlining Databases: Supporting Fast Transactional Workloads on Universal Columnar Data File FormatsTianyu Li, Matthew Butrovich, Amadou Ngom, Wan Shen Lim 等VLDB 2021 · 被引用 28 次
- Nested Parquet Is Flat, Why Not Use It? How To Scan Nested Data With On-the-Fly Key Generation and JoinsAlice Rey, Maximilian Rieger, Thomas NeumannSIGMOD 2025 · 被引用 1 次
- LST-Bench: Benchmarking Log-Structured Tables in the CloudJesús Camacho-Rodríguez, Ashvin Agrawal, Anja Gruenheid, Ashit Gosalia 等SIGMOD 2024 · 被引用 7 次
- The FastLanes Compression Layout: Decoding >100 Billion Integers per Second with Scalar CodeAzim Afroozeh, Peter BonczVLDB 2023 · 被引用 44 次
