F3: The Open-Source Data File Format for the Future
Xinyu Zeng, Ruijun Meng, Martin Prammer, Wes McKinney, Jignesh M. Patel, Andrew Pavlo, Huanchen Zhang
摘要
Columnar storage formats are the foundation for modern data analytics systems. The proliferation of open-source file formats (i.e., Parquet, ORC) allows seamless data sharing across disparate platforms. However, these formats were created over a decade ago for hardware and workload environments that are much different from today. Although these formats have incorporated some updates to their specification to adapt to these changes, not all deployments support those modifications, and too often systems cannot overcome the formats' deficiencies and limitations without a rewrite. In this paper, we present the F uture-proof File Format (F3) project. It is a next-generation open-source file format with interoperability, extensibility, and efficiency as its core design principles. F3 obviates the need to create a new format every time a shift occurs in data processing and computing by providing a data organization structure and a general-purpose API to allow developers to add new encoding schemes easily. Each self-describing F3 file includes both the data and meta-data, as well as WebAssembly (Wasm) binaries to decode the data. Embedding the decoders in each file requires minimal storage (kilobytes) and ensures compatibility on any platform in case native decoders are unavailable. To evaluate F3, we compared it against legacy and state-of-the-art open-source file formats. Our evaluations demonstrate the efficacy of F3's storage layout and the benefits of Wasm-driven decoding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Active Data Lakes: Regaining Physical Data Independence Without Losing InteroperabilityPascal Ginter, Viktor LeisVLDB 2026 · 被引用 4 次
- BtrLog: Low-Latency Logging for Cloud Database SystemsMaximilian Kuschewski, Lam-Duy Nguyen, Matthias Jasny, Tobias Ziegler 等VLDB 2026
- LiquidCache: Efficient Pushdown Caching for Cloud-Native Data AnalyticsXiangpeng Hao, Andrew Lamb, Yibo Wu, Andrea C. Arpaci-Dusseau 等VLDB 2025
它引用的顶会 Paper9
- Chimp: Efficient Lossless Floating Point Compression for Time Series DatabasesPanagiotis Liakos, Katia Papakonstantinopoulou, Yannis KotidisVLDB 2022 · 被引用 76 次
- An Empirical Evaluation of Columnar Storage FormatsXinyu Zeng, Yulong Hui, Jiahong Shen, Andrew Pavlo 等VLDB 2024 · 被引用 59 次
- BtrBlocks: Efficient Columnar Compression for Data LakesMaximilian Kuschewski, David Sauerwein, Adnan Alhomssi, Viktor LeisSIGMOD 2023 · 被引用 47 次
- Exploiting Cloud Object Storage for High-Performance AnalyticsDominik Durner, Viktor Leis, Thomas NeumannVLDB 2023 · 被引用 45 次
- The FastLanes Compression Layout: Decoding >100 Billion Integers per Second with Scalar CodeAzim Afroozeh, Peter BonczVLDB 2023 · 被引用 44 次
相关 Paper
- AnyBlox: A Framework for Self-Decoding DatasetsMateusz Gienieczko, Maximilian Kuschewski, Thomas Neumann, Viktor Leis 等VLDB 2025 · 被引用 3 次
- A Deep Dive into Common Open Formats for Analytical DBMSsChunwei Liu, Anna Pavlenko, Matteo Interlandi, Brandon HaynesVLDB 2023 · 被引用 23 次
- Selection Pushdown in Column Stores using Bit Manipulation InstructionsYinan Li, Jianan Lu, Badrish ChandramouliSIGMOD 2023 · 被引用 15 次
- Nested Parquet Is Flat, Why Not Use It? How To Scan Nested Data With On-the-Fly Key Generation and JoinsAlice Rey, Maximilian Rieger, Thomas NeumannSIGMOD 2025 · 被引用 1 次
- Multi-modal Learning for WebAssembly Reverse EngineeringHanxian Huang, Jishen ZhaoISSTA 2024 · 被引用 1 次
