Pando: Enhanced Data Skipping with Logical Data Partitioning
Sivaprasad Sudhir, Wenbo Tao, Nikolay Pavlovich Laptev, Cyrille Habis, Michael J. Cafarella, Samuel Madden
摘要
With enormous volumes of data, quickly retrieving data that is relevant to a query is essential for achieving high performance. Modern cloud-based database systems often partition the data into blocks and employ various techniques to skip irrelevant blocks during query execution. Several algorithms, often based on historical properties of a workload of queries run over the data, have been proposed to tune the physical layout of data to reduce the number of blocks accessed. The effectiveness of these methods at skipping blocks depends on what metadata is stored and how well the physical data layout aligns with the queries. Existing work on automatic physical database design misses significant opportunities in skipping blocks because it ignores logical predicates in the workload that exhibit strongly correlated results. In this paper, we present Pando which enables significantly better block skipping than past methods by informing physical layout decisions with correlation-aware logical partitioning. Across a range of benchmark and real-world workloads, Pando attains up to 2.8X reduction in the number of blocks scanned and up to 2.3X speedup in end-to-end query execution time over the state-of-the-art techniques.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- HONEYBEE: Efficient Role-based Access Control for Vector Databases via Dynamic PartitioningHongbin Zhong, Matthew Lentz, Nina Narodytska, Adriana Szekeres 等SIGMOD 2026 · 被引用 5 次
- Optimizing Collections of Bloom Filters within a Space BudgetGabriel Mersy, Zhuo Wang, Stavros Sintos, Sanjay KrishnanVLDB 2024 · 被引用 2 次
- Partition, Don't Sort! Compression Boosters for Cloud Data Ingestion PipelinesPatrick Hansert, Sebastian MichelVLDB 2024 · 被引用 2 次
- Dynamic Data Layout Optimization with Worst-Case GuaranteesKexin Rong, Paul Liu, Sarah Ashok Sonje, Moses CharikarICDE 2024 · 被引用 1 次
- Bonsai: Compiling Queries to Pruned Tree TraversalsAlexander J. Root, Christophe Gyurgyik, Purvi Goel, Kayvon Fatahalian 等PLDI 2026
它引用的顶会 Paper6
- Tsunami: A Learned Multi-dimensional Index for Correlated Data and Skewed WorkloadsJialin Ding, Vikram Nathan, Mohammad Alizadeh, Tim KraskaVLDB 2021 · 被引用 178 次
- Qd-tree: Learning Data Layouts for Big Data AnalyticsZongheng Yang, Badrish Chandramouli, Chi Wang, Johannes Gehrke 等SIGMOD 2020 · 被引用 87 次
- Learning a Partitioning Advisor for Cloud DatabasesBenjamin Hilprecht, Carsten Binnig, Uwe RöhmSIGMOD 2020 · 被引用 64 次
- Instance-Optimized Data Layouts for Cloud Analytics WorkloadsJialin Ding, Umar Farooq Minhas, Badrish Chandramouli, Chi Wang 等SIGMOD 2021 · 被引用 37 次
- Pushing Data-Induced Predicates Through Joins in Big-Data ClustersLaurel J. Orr, Srikanth Kandula, Surajit ChaudhuriVLDB 2020 · 被引用 35 次
相关 Paper
- Provenance-based Data SkippingXing Niu, Boris Glavic, Ziyu Liu, Pengyuan Li 等VLDB 2022 · 被引用 9 次
- Sieve: A Learned Data-Skipping Index for Data AnalyticsYulai Tong, Jiazhen Liu, Hua Wang, Ke Zhou 等VLDB 2023 · 被引用 10 次
- Data Chunk Compaction in Vectorized ExecutionYiming Qiao, Huanchen ZhangSIGMOD 2025 · 被引用 2 次
- Deductive optimization of relational data storageJohn K. Feser, Sam Madden, Nan Tang, Armando Solar-LezamaOOPSLA 2020 · 被引用 5 次
- An Efficient Transfer Learning Based Configuration Adviser for Database TuningXinyi Zhang, Hong Wu, Yang Li, Zhengju Tang 等VLDB 2024 · 被引用 25 次
