Fusion: An Analytics Object Store Optimized for Query Pushdown
Jianan Lu, Ashwini Raina, Asaf Cidon, Michael J. Freedman
Abstract
The prevalence of disaggregated storage in public clouds has led to increased latency in modern OLAP cloud databases, particularly when handling ad-hoc and highly-selective queries on large objects. To address this, cloud databases have adopted computation pushdown, executing query predicates closer to the storage layer. However, existing pushdown solutions are inefficient in erasure-coded storage. Cloud storage employs erasure coding that partitions analytics file objects into fixed-sized blocks and distributes them across storage nodes. Consequently, when a specific part of the object is queried, the storage system must reassemble the object across nodes, incurring significant network latency.
In this work, we present Fusion, an object store for analytics that is optimized for query pushdown on erasurecoded data. It co-designs its erasure coding and file placement topologies, taking into account popular analytics file formats (e.g., Parquet). Fusion employs a novel stripe construction algorithm that prevents fragmentation of computable units within an object, and minimizes storage overhead during erasure coding. Compared to existing erasure-coded stores, Fusion improves median and tail latency by 64% and 81%, respectively, on TPC-H, and up to 40% and 48% respectively, on real-world SQL queries. Fusion achieves this while incurring a modest 1.2% storage overhead compared to the optimal.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e811281b-d867-4aa0-91ef-2abaef6770b2Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Exploiting Combined Locality for Wide-Stripe Erasure Coding in Distributed StorageYuchong Hu, Liangfeng Cheng, Qiaori Yao, Patrick P. C. Lee et al.FAST 2021 · 88 citations
- FlexPushdownDB: Hybrid Pushdown and Caching in a Cloud DBMSYifei Yang, Matt Youill, Matthew E. Woicik, Yizhou Liu et al.VLDB 2021 · 67 citations
- An Empirical Evaluation of Columnar Storage FormatsXinyu Zeng, Yulong Hui, Jiahong Shen, Andrew Pavlo et al.VLDB 2024 · 59 citations
- Ship Compute or Ship Data? Why Not Both?Jie You, Jingfeng Wu, Xin Jin, Mosharaf ChowdhuryNSDI 2021 · 25 citations
- Adaptive Placement for In-memory Storage FunctionsAnkit Bhardwaj, Chinmay Kulkarni, Ryan StutsmanUSENIX ATC 2020 · 18 citations
Related papers
- Crystal: A Unified Cache Storage System for Analytical DatabasesDominik Durner, Badrish Chandramouli, Yinan LiVLDB 2021 · 11 citations
- Exploiting Cloud Object Storage for High-Performance AnalyticsDominik Durner, Viktor Leis, Thomas NeumannVLDB 2023 · 45 citations
- Generalized Sub-Query Fusion for Eliminating Redundant I/O from Big-Data QueriesPartho Sarthi, Kaushik Rajan, Akash Lal, Abhishek Modi et al.OSDI 2020 · 5 citations
- LEGOStore: A Linearizable Geo-Distributed Store Combining Replication and Erasure CodingHamidReza Zare, Viveck R. Cadambe, Bhuvan Urgaonkar, Nader Alfares et al.VLDB 2022 · 10 citations
- BtrBlocks: Efficient Columnar Compression for Data LakesMaximilian Kuschewski, David Sauerwein, Adnan Alhomssi, Viktor LeisSIGMOD 2023 · 47 citations
