The FastLanes File Format
Azim Afroozeh, Peter Boncz
摘要
This paper introduces a new open-source big data file format, called FastLanes. It is designed for modern data-parallel execution (SIMD or GPU), and evolves the features of previous data formats such as Parquet, which are the foundation of data lakes, and which increasingly are used in AI pipelines. It does so by avoiding generic compression methods (e.g. Snappy) in favor of lightweight encodings, that are fully data-parallel. To enhance compression ratio, it cascades encodings using a flexible expression encoding mechanism. This mechanism also enables multi-column compression (MCC), enhancing compression by exploiting correlations between columns, a long-time weakness of columnar storage. We contribute a 2-phase algorithm to find encodings expressions during compression. FastLanes also innovates in its API, providing flexible support for partial decompression, facilitating engines to execute queries on compressed data. FastLanes is designed for fine-grained access, at the level of small batches rather than rowgroups; in order to limit the decompression memory footprint to fit CPU and GPU caches. We contribute an open-source implementation of FastLanes in portable (auto-vectorizing) C++. Our evaluation on a corpus of real-world data shows that FastLanes improves compression ratio over Parquet, while strongly accelerating decompression, making it a win-win over the state-of-the-art.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Active Data Lakes: Regaining Physical Data Independence Without Losing InteroperabilityPascal Ginter, Viktor LeisVLDB 2026 · 被引用 4 次
- L3: A GPU-Native Co-Designed Data Format for Learned Lossless Lightweight CompressionYouyang Xia, Feng Zhang, Junda Pan, Yihao Liu 等SIGMOD 2026 · 被引用 1 次
- Accelerating String-Heavy Queries with LLM Token TablesTobias Schmidt, Nicolas Schmitt, Thomas Neumann, Andreas KipfVLDB 2026
- LiquidCache: Efficient Pushdown Caching for Cloud-Native Data AnalyticsXiangpeng Hao, Andrew Lamb, Yibo Wu, Andrea C. Arpaci-Dusseau 等VLDB 2025
它引用的顶会 Paper16
- An Empirical Evaluation of Columnar Storage FormatsXinyu Zeng, Yulong Hui, Jiahong Shen, Andrew Pavlo 等VLDB 2024 · 被引用 59 次
- BtrBlocks: Efficient Columnar Compression for Data LakesMaximilian Kuschewski, David Sauerwein, Adnan Alhomssi, Viktor LeisSIGMOD 2023 · 被引用 47 次
- CompressDB: Enabling Efficient Compressed Data Direct Processing for Various DatabasesFeng Zhang, Weitao Wan, Chenyang Zhang, Jidong Zhai 等SIGMOD 2022 · 被引用 46 次
- Tile-based Lightweight Integer Compression in GPUAnil Shanbhag, Bobbi W. Yogatama, Xiangyao Yu, Samuel MaddenSIGMOD 2022 · 被引用 45 次
- Good to the Last Bit: Data-Driven Encoding with CodecDBHao Jiang, Chunwei Liu, John Paparrizos, Andrew A. Chien 等SIGMOD 2021 · 被引用 45 次
相关 Paper
- The FastLanes Compression Layout: Decoding >100 Billion Integers per Second with Scalar CodeAzim Afroozeh, Peter BonczVLDB 2023 · 被引用 44 次
- GPU Acceleration of SQL Analytics on Compressed DataZezhou Huang, Krystian Sakowski, Hans Lehnert, Wei Cui 等VLDB 2026 · 被引用 1 次
- ALP: Adaptive Lossless floating-Point CompressionAzim Afroozeh, Leonardo Kuffó, Peter BonczSIGMOD 2024 · 被引用 33 次
- Efficient Lossless Compression of Scientific Floating-Point Data on CPUs and GPUsNoushin Azami, Alex Fallin, Martin BurtscherASPLOS 2025 · 被引用 18 次
- Optimizing Random Access to Hierarchically-Compressed Data on GPUFeng Zhang, Yihua Hu, Haipeng Ding, Zhiming Yao 等SC 2022 · 被引用 5 次
