FlatStor: An Efficient Embedded-Index Based Columnar Data Layout for Multimodal Data Workloads
Chi Zhang, Shihao Zhang, Yunfei Gu, Chentao Wu, Jie Li, Qin Zhang, Xusheng Chen, Jie Meng
Abstract
Modern data lakes have become essential for storing, managing, and analyzing massive amounts of heterogeneous data. As production data increasingly exhibits multimodal storage characteristics and multi-purpose access patterns, efficient management of such complexities becomes critical. However, current hybrid storage system-based data lakes face persistent challenges, including synchronization overhead, data correlation disruption, and escalating storage costs due to the involvement of multiple underlying storage systems. While columnar storage, central to data lakes, addresses hybrid-system inefficiencies, it struggles with the complexities of multimodal data storage and multi-purpose access. To tackle these challenges, we analyze access patterns across various scenarios and assess the issues in storing multimodal data. Based on these insights, we propose FlatStor, a FlatBuffers-based columnar Storage format with embedded indexing. It supports point access through indexing and handles multimodal data by vertically partitioning and treating each modality as a byte stream for storage. It also applies FSST compression, reducing storage overhead significantly. Benchmark evaluations reveal that FlatStor reduces the access latency by 99.6% and the storage overhead by 91.3% compared to Parquet in inference workloads. Furthermore, FlatStor outperforms LanceV2 with a 41.3% latency improvement, maintaining minimal additional overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3e86ecc-8c22-4116-b0d1-9a0bb7dd69aeBuilds on8
- On Uni-Modal Feature Learning in Supervised Multi-Modal LearningChenzhuang Du, Jiaye Teng, Tingle Li, Yichen Liu et al.ICML 2023 · 79 citations
- Clairvoyant prefetching for distributed machine learning I/ONikoli Dryden, Roman Böhringer, Tal Ben-Nun, Torsten HoeflerSC 2021 · 60 citations
- An Empirical Evaluation of Columnar Storage FormatsXinyu Zeng, Yulong Hui, Jiahong Shen, Andrew Pavlo et al.VLDB 2024 · 59 citations
- BtrBlocks: Efficient Columnar Compression for Data LakesMaximilian Kuschewski, David Sauerwein, Adnan Alhomssi, Viktor LeisSIGMOD 2023 · 47 citations
- SHADE: Enable Fundamental Cacheability for Distributed Deep Learning TrainingRedwan Ibne Seraj Khan, Ahmad Hossein Yazdani, Yuqi Fu, Arnab K. Paul et al.FAST 2023 · 29 citations
Related papers
- GraphAr: An Efficient Storage Scheme for Graph Data in Data LakesXue Li, Weibin Zeng, Zhibin Wang, Diwen Zhu et al.VLDB 2025 · 5 citations
- Rottnest: Indexing Data Lakes for SearchZiheng Wang, Sasha Krassovsky, Conor Kennedy, Alex Aiken et al.ICDE 2025 · 1 citation
- ARCADE: A Real-Time Data System for Hybrid and Continuous Query Processing Across Diverse Data ModalitiesJingyi Yang, Songsong Mo, Jiachen Shi, Zihao Yu et al.ICDE 2026 · 1 citation
- Active Data Lakes: Regaining Physical Data Independence Without Losing InteroperabilityPascal Ginter, Viktor LeisVLDB 2026 · 4 citations
- LST-Bench: Benchmarking Log-Structured Tables in the CloudJesús Camacho-Rodríguez, Ashvin Agrawal, Anja Gruenheid, Ashit Gosalia et al.SIGMOD 2024 · 7 citations
