FlatStor: An Efficient Embedded-Index Based Columnar Data Layout for Multimodal Data Workloads
Chi Zhang, Shihao Zhang, Yunfei Gu, Chentao Wu, Jie Li, Qin Zhang, Xusheng Chen, Jie Meng
摘要
Modern data lakes have become essential for storing, managing, and analyzing massive amounts of heterogeneous data. As production data increasingly exhibits multimodal storage characteristics and multi-purpose access patterns, efficient management of such complexities becomes critical. However, current hybrid storage system-based data lakes face persistent challenges, including synchronization overhead, data correlation disruption, and escalating storage costs due to the involvement of multiple underlying storage systems. While columnar storage, central to data lakes, addresses hybrid-system inefficiencies, it struggles with the complexities of multimodal data storage and multi-purpose access. To tackle these challenges, we analyze access patterns across various scenarios and assess the issues in storing multimodal data. Based on these insights, we propose FlatStor, a FlatBuffers-based columnar Storage format with embedded indexing. It supports point access through indexing and handles multimodal data by vertically partitioning and treating each modality as a byte stream for storage. It also applies FSST compression, reducing storage overhead significantly. Benchmark evaluations reveal that FlatStor reduces the access latency by 99.6% and the storage overhead by 91.3% compared to Parquet in inference workloads. Furthermore, FlatStor outperforms LanceV2 with a 41.3% latency improvement, maintaining minimal additional overhead.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- On Uni-Modal Feature Learning in Supervised Multi-Modal LearningChenzhuang Du, Jiaye Teng, Tingle Li, Yichen Liu 等ICML 2023 · 被引用 79 次
- Clairvoyant prefetching for distributed machine learning I/ONikoli Dryden, Roman Böhringer, Tal Ben-Nun, Torsten HoeflerSC 2021 · 被引用 60 次
- An Empirical Evaluation of Columnar Storage FormatsXinyu Zeng, Yulong Hui, Jiahong Shen, Andrew Pavlo 等VLDB 2024 · 被引用 59 次
- BtrBlocks: Efficient Columnar Compression for Data LakesMaximilian Kuschewski, David Sauerwein, Adnan Alhomssi, Viktor LeisSIGMOD 2023 · 被引用 47 次
- SHADE: Enable Fundamental Cacheability for Distributed Deep Learning TrainingRedwan Ibne Seraj Khan, Ahmad Hossein Yazdani, Yuqi Fu, Arnab K. Paul 等FAST 2023 · 被引用 29 次
相关 Paper
- GraphAr: An Efficient Storage Scheme for Graph Data in Data LakesXue Li, Weibin Zeng, Zhibin Wang, Diwen Zhu 等VLDB 2025 · 被引用 5 次
- Rottnest: Indexing Data Lakes for SearchZiheng Wang, Sasha Krassovsky, Conor Kennedy, Alex Aiken 等ICDE 2025 · 被引用 1 次
- ARCADE: A Real-Time Data System for Hybrid and Continuous Query Processing Across Diverse Data ModalitiesJingyi Yang, Songsong Mo, Jiachen Shi, Zihao Yu 等ICDE 2026 · 被引用 1 次
- Active Data Lakes: Regaining Physical Data Independence Without Losing InteroperabilityPascal Ginter, Viktor LeisVLDB 2026 · 被引用 4 次
- LST-Bench: Benchmarking Log-Structured Tables in the CloudJesús Camacho-Rodríguez, Ashvin Agrawal, Anja Gruenheid, Ashit Gosalia 等SIGMOD 2024 · 被引用 7 次
