Sampling-based Predictive Database Buffer Management
Theo Vanderkooy, Mohammad Khalaji, Runsheng Benson Guo, Khuzaima Daudjee
Abstract
Systems often need to support analytical (OLAP) workloads that perform concurrent scans of data on secondary storage. The buffer manager is tasked with fetching data into the database system's buffer pool and caching it there so as to increase the hit rate on these data, thereby lowering query latencies. This paper presents a database buffer caching policy that uses information about long-running scans to estimate future accesses. These estimates are used to approximate an optimal buffer caching policy that would otherwise infeasibly require knowledge about future accesses. Since a buffer caching policy must be efficient with low overhead, we present sampling-based predictive buffer management techniques where buffer eviction considers only a small random sample of buffers and access time estimates are used to select from the sample. This design is advantageous as it is easily tuned by adjusting the sample size, and easily modified to improve access time estimates and to expand the set of workload types that can be predicted effectively. We evaluate our techniques through both simulation studies on real Amazon Redshift workload traces and through implementation into the well-known open-source PostgreSQL database system on the popular TPC-H and YCSB benchmarks. We show that our approach delivers substantial performance improvements for workloads with scans, reducing I/O volume significantly by up to 40% over Post-greSQL's Clock-sweep policy and over prior predictive approaches for workloads using sequential scans and index accesses.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 62dff7d0-4612-4609-bd8e-64a7dec380beBuilds on4
- Learning Relaxed Belady for Content Distribution Network CachingZhenyu Song, Daniel S. Berger, Kai Li, Wyatt LloydNSDI 2020 · 193 citations
- Why TPC Is Not Enough: An Analysis of the Amazon Redshift FleetAlexander van Renen, Dominik Horn, Pascal Pfeil, Kapil Vaidya et al.VLDB 2024 · 65 citations
- Effective Mimicry of Belady's MIN PolicyIshan Shah, Akanksha Jain, Calvin LinHPCA 2022 · 44 citations
- Write-Aware Timestamp Tracking: Effective and Efficient Page Replacement for Modern HardwareDemian E. Vöhringer, Viktor LeisVLDB 2023 · 11 citations
Related papers
- ACEing the Bufferpool Management Paradigm for Modern Storage DevicesTarikul Islam Papon, Manos AthanassoulisICDE 2023 · 8 citations
- High-Performance DBMSs with io_uring: When and How to Use ItMatthias Jasny, Muhammad El-Hindi, Tobias Ziegler, Viktor Leis et al.VLDB 2026 · 5 citations
- Generalizable Address-Aware Semantic Prefetching for Scalable Transactional and Analytical WorkloadsFarzaneh Zirak, Farhana Choudhury, Renata Borovica-GajicICDE 2026 · 1 citation
- LBSC: A Cost-Aware Caching Framework for Cloud DatabasesZhaoxuan Ji, Zhongle Xie, Yuncheng Wu, Meihui ZhangICDE 2024 · 7 citations
- Predictive Translation: High-Performance Buffer Management Without the Trade-OffsMichael Zinsmeister, Lam-Duy Nguyen, Viktor Leis, Thomas NeumannSIGMOD 2026 · 3 citations
