Crystal: A Unified Cache Storage System for Analytical Databases
Dominik Durner, Badrish Chandramouli, Yinan Li
Abstract
Cloud analytical databases employ a disaggregated storage model, where the elastic compute layer accesses data persisted on remote cloud storage in block-oriented columnar formats. Given the high latency and low bandwidth to remote storage and the limited size of fast local storage, caching data at the compute node is important and has resulted in a renewed interest in caching for analytics. Today, each DBMS builds its own caching solution, usually based on file-or block-level LRU. In this paper, we advocate a new architecture of a smart cache storage system called Crystal , that is co-located with compute. Crystal's clients are DBMS-specific "data sources" with push-down predicates. Similar in spirit to a DBMS, Crystal incorporates query processing and optimization components focusing on efficient caching and serving of single-table hyper-rectangles called regions. Results show that Crystal, with a small DBMS-specific data source connector, can significantly improve query latencies on unmodified Spark and Greenplum while also saving on bandwidth from remote storage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7e259679-2612-4499-8d71-c85fdb28bdc6Cited by top-tier papers8
- ScaleStore: A Fast and Cost-Efficient Storage Engine using DRAM, NVMe, and RDMATobias Ziegler, Carsten Binnig, Viktor LeisSIGMOD 2022 · 51 citations
- Predicate Pushdown for Data Science PipelinesCong Yan, Yin Lin, Yeye HeSIGMOD 2023 · 15 citations
- Selection Pushdown in Column Stores using Bit Manipulation InstructionsYinan Li, Jianan Lu, Badrish ChandramouliSIGMOD 2023 · 15 citations
- On-Demand State Separation for Cloud Data WarehousingChristian Winter, Jana Giceva, Thomas Neumann, Alfons KemperVLDB 2022 · 9 citations
- uCache: A Customizable Unikernel-based IO CacheIlya Meignan-Masson, Masanori Misono, Viktor Leis, Pramod BhatotiaFAST 2026 · 1 citation
Builds on2
- Tsunami: A Learned Multi-dimensional Index for Correlated Data and Skewed WorkloadsJialin Ding, Vikram Nathan, Mohammad Alizadeh, Tim KraskaVLDB 2021 · 178 citations
- Pushing Data-Induced Predicates Through Joins in Big-Data ClustersLaurel J. Orr, Srikanth Kandula, Surajit ChaudhuriVLDB 2020 · 35 citations
Related papers
- Exploiting Cloud Object Storage for High-Performance AnalyticsDominik Durner, Viktor Leis, Thomas NeumannVLDB 2023 · 45 citations
- FlexPushdownDB: Hybrid Pushdown and Caching in a Cloud DBMSYifei Yang, Matt Youill, Matthew E. Woicik, Yizhou Liu et al.VLDB 2021 · 67 citations
- Fusion: An Analytics Object Store Optimized for Query PushdownJianan Lu, Ashwini Raina, Asaf Cidon, Michael J. FreedmanASPLOS 2025 · 4 citations
- Persistent Memory Disaggregation for Cloud-Native Relational DatabasesChaoyi Ruan, Yingqiang Zhang, Chao Bi, Xiaosong Ma et al.ASPLOS 2023 · 24 citations
- A Community Cache with Complete InformationMania Abdi, Amin Mosayyebzadeh, Mohammad Hossein Hajkazemi, Emine Ugur Kaynar et al.FAST 2021 · 2 citations
