F3: The Open-Source Data File Format for the Future
Xinyu Zeng, Ruijun Meng, Martin Prammer, Wes McKinney, Jignesh M. Patel, Andrew Pavlo, Huanchen Zhang
Abstract
Columnar storage formats are the foundation for modern data analytics systems. The proliferation of open-source file formats (i.e., Parquet, ORC) allows seamless data sharing across disparate platforms. However, these formats were created over a decade ago for hardware and workload environments that are much different from today. Although these formats have incorporated some updates to their specification to adapt to these changes, not all deployments support those modifications, and too often systems cannot overcome the formats' deficiencies and limitations without a rewrite. In this paper, we present the F uture-proof File Format (F3) project. It is a next-generation open-source file format with interoperability, extensibility, and efficiency as its core design principles. F3 obviates the need to create a new format every time a shift occurs in data processing and computing by providing a data organization structure and a general-purpose API to allow developers to add new encoding schemes easily. Each self-describing F3 file includes both the data and meta-data, as well as WebAssembly (Wasm) binaries to decode the data. Embedding the decoders in each file requires minimal storage (kilobytes) and ensures compatibility on any platform in case native decoders are unavailable. To evaluate F3, we compared it against legacy and state-of-the-art open-source file formats. Our evaluations demonstrate the efficacy of F3's storage layout and the benefits of Wasm-driven decoding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 956789ff-7f0e-4bf5-8893-65b730962662Cited by top-tier papers3
- Active Data Lakes: Regaining Physical Data Independence Without Losing InteroperabilityPascal Ginter, Viktor LeisVLDB 2026 · 4 citations
- BtrLog: Low-Latency Logging for Cloud Database SystemsMaximilian Kuschewski, Lam-Duy Nguyen, Matthias Jasny, Tobias Ziegler et al.VLDB 2026
- LiquidCache: Efficient Pushdown Caching for Cloud-Native Data AnalyticsXiangpeng Hao, Andrew Lamb, Yibo Wu, Andrea C. Arpaci-Dusseau et al.VLDB 2025
Builds on9
- Chimp: Efficient Lossless Floating Point Compression for Time Series DatabasesPanagiotis Liakos, Katia Papakonstantinopoulou, Yannis KotidisVLDB 2022 · 76 citations
- An Empirical Evaluation of Columnar Storage FormatsXinyu Zeng, Yulong Hui, Jiahong Shen, Andrew Pavlo et al.VLDB 2024 · 59 citations
- BtrBlocks: Efficient Columnar Compression for Data LakesMaximilian Kuschewski, David Sauerwein, Adnan Alhomssi, Viktor LeisSIGMOD 2023 · 47 citations
- Exploiting Cloud Object Storage for High-Performance AnalyticsDominik Durner, Viktor Leis, Thomas NeumannVLDB 2023 · 45 citations
- The FastLanes Compression Layout: Decoding >100 Billion Integers per Second with Scalar CodeAzim Afroozeh, Peter BonczVLDB 2023 · 44 citations
Related papers
- AnyBlox: A Framework for Self-Decoding DatasetsMateusz Gienieczko, Maximilian Kuschewski, Thomas Neumann, Viktor Leis et al.VLDB 2025 · 3 citations
- A Deep Dive into Common Open Formats for Analytical DBMSsChunwei Liu, Anna Pavlenko, Matteo Interlandi, Brandon HaynesVLDB 2023 · 23 citations
- Selection Pushdown in Column Stores using Bit Manipulation InstructionsYinan Li, Jianan Lu, Badrish ChandramouliSIGMOD 2023 · 15 citations
- Nested Parquet Is Flat, Why Not Use It? How To Scan Nested Data With On-the-Fly Key Generation and JoinsAlice Rey, Maximilian Rieger, Thomas NeumannSIGMOD 2025 · 1 citation
- Multi-modal Learning for WebAssembly Reverse EngineeringHanxian Huang, Jishen ZhaoISSTA 2024 · 1 citation
