Pando: Enhanced Data Skipping with Logical Data Partitioning
Sivaprasad Sudhir, Wenbo Tao, Nikolay Pavlovich Laptev, Cyrille Habis, Michael J. Cafarella, Samuel Madden
Abstract
With enormous volumes of data, quickly retrieving data that is relevant to a query is essential for achieving high performance. Modern cloud-based database systems often partition the data into blocks and employ various techniques to skip irrelevant blocks during query execution. Several algorithms, often based on historical properties of a workload of queries run over the data, have been proposed to tune the physical layout of data to reduce the number of blocks accessed. The effectiveness of these methods at skipping blocks depends on what metadata is stored and how well the physical data layout aligns with the queries. Existing work on automatic physical database design misses significant opportunities in skipping blocks because it ignores logical predicates in the workload that exhibit strongly correlated results. In this paper, we present Pando which enables significantly better block skipping than past methods by informing physical layout decisions with correlation-aware logical partitioning. Across a range of benchmark and real-world workloads, Pando attains up to 2.8X reduction in the number of blocks scanned and up to 2.3X speedup in end-to-end query execution time over the state-of-the-art techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca819d5e-3799-4b1d-a5ea-f2a5e339f777Cited by top-tier papers6
- HONEYBEE: Efficient Role-based Access Control for Vector Databases via Dynamic PartitioningHongbin Zhong, Matthew Lentz, Nina Narodytska, Adriana Szekeres et al.SIGMOD 2026 · 5 citations
- Optimizing Collections of Bloom Filters within a Space BudgetGabriel Mersy, Zhuo Wang, Stavros Sintos, Sanjay KrishnanVLDB 2024 · 2 citations
- Partition, Don't Sort! Compression Boosters for Cloud Data Ingestion PipelinesPatrick Hansert, Sebastian MichelVLDB 2024 · 2 citations
- Dynamic Data Layout Optimization with Worst-Case GuaranteesKexin Rong, Paul Liu, Sarah Ashok Sonje, Moses CharikarICDE 2024 · 1 citation
- Bonsai: Compiling Queries to Pruned Tree TraversalsAlexander J. Root, Christophe Gyurgyik, Purvi Goel, Kayvon Fatahalian et al.PLDI 2026
Builds on6
- Tsunami: A Learned Multi-dimensional Index for Correlated Data and Skewed WorkloadsJialin Ding, Vikram Nathan, Mohammad Alizadeh, Tim KraskaVLDB 2021 · 178 citations
- Qd-tree: Learning Data Layouts for Big Data AnalyticsZongheng Yang, Badrish Chandramouli, Chi Wang, Johannes Gehrke et al.SIGMOD 2020 · 87 citations
- Learning a Partitioning Advisor for Cloud DatabasesBenjamin Hilprecht, Carsten Binnig, Uwe RöhmSIGMOD 2020 · 64 citations
- Instance-Optimized Data Layouts for Cloud Analytics WorkloadsJialin Ding, Umar Farooq Minhas, Badrish Chandramouli, Chi Wang et al.SIGMOD 2021 · 37 citations
- Pushing Data-Induced Predicates Through Joins in Big-Data ClustersLaurel J. Orr, Srikanth Kandula, Surajit ChaudhuriVLDB 2020 · 35 citations
Related papers
- Provenance-based Data SkippingXing Niu, Boris Glavic, Ziyu Liu, Pengyuan Li et al.VLDB 2022 · 9 citations
- Sieve: A Learned Data-Skipping Index for Data AnalyticsYulai Tong, Jiazhen Liu, Hua Wang, Ke Zhou et al.VLDB 2023 · 10 citations
- Data Chunk Compaction in Vectorized ExecutionYiming Qiao, Huanchen ZhangSIGMOD 2025 · 2 citations
- Deductive optimization of relational data storageJohn K. Feser, Sam Madden, Nan Tang, Armando Solar-LezamaOOPSLA 2020 · 5 citations
- An Efficient Transfer Learning Based Configuration Adviser for Database TuningXinyi Zhang, Hong Wu, Yang Li, Zhengju Tang et al.VLDB 2024 · 25 citations
