A Community Cache with Complete Information
Mania Abdi, Amin Mosayyebzadeh, Mohammad Hossein Hajkazemi, Emine Ugur Kaynar, Ata Turk, Larry Rudolph, Orran Krieger, Peter Desnoyers
Abstract
Kariz is a new architecture for caching data from datalakes accessed, potentially concurrently, by multiple analytic platforms. It integrates rich information from analytics platforms with global knowledge about demand and resource availability to enable sophisticated cache management and prefetching strategies that, for example, combine historical run time information with job dependency graphs (DAGs), information about the cache state and sharing across compute clusters. Our prototype supports multiple analytic frameworks (Pig/Hadoop and Spark), and we show that the required changes are modest. We have implemented three algorithms in Kariz for optimizing the caching of individual queries (one from the literature, and two novel to our platform) and three policies for optimizing across queries from, potentially, multiple different clusters. With an algorithm that fully exploits the rich information available from Kariz, we demonstrate major speedups (as much as 3×) for TPC-H and TPC-DS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 82c52463-4e38-4140-b006-b7accc9c781dCited by top-tier papers2
- Adaptive Online Cache Capacity Optimization via Lightweight Working Set Size Estimation at ScaleRong Gu, Simian Li, Haipeng Dai, Hancheng Wang et al.USENIX ATC 2023 · 17 citations
- ScalaCache: Scalable User-Space Page Cache Management with Software-Hardware CoordinationLi Peng, Yuda An, You Zhou, Chenxi Wang et al.USENIX ATC 2024 · 6 citations
Builds on2
Related papers
- Whiz: Data-Driven Analytics ExecutionRobert Grandl, Arjun Singhvi, Raajay Viswanathan, Aditya AkellaNSDI 2021 · 9 citations
- Cache-Efficient Top-k Aggregation over High Cardinality Large DatasetsTarique Siddiqui, Vivek R. Narasayya, Marius Dumitru, Surajit ChaudhuriVLDB 2024
- PTO: A Workload-driven Predictive Table Optimizer for Lakehouse SystemsVenkata Vamsikrishna Meduri, David Kreismann, Ronald Barber, Berthold ReinwaldSIGMOD 2026
- Crystal: A Unified Cache Storage System for Analytical DatabasesDominik Durner, Badrish Chandramouli, Yinan LiVLDB 2021 · 11 citations
- LakeHelm: Zero-Shot Lakehouse Advisor for Joint Engine-Format Selection and ConfigurationZhongwei Xu, Siyuan Dong, Haotian Gong, Donna Pham et al.VLDB 2026
