Sieve: A Learned Data-Skipping Index for Data Analytics
Yulai Tong, Jiazhen Liu, Hua Wang, Ke Zhou, Rongfeng He, Qin Zhang, Cheng Wang
摘要
Modern data analytics services are coupled with external data storage services, making I/O from remote cloud storage one of the dominant costs for query processing. Techniques such as columnar block-based data organization and compression have become standard practices for these services to save storage and processing cost. However, the problem of effectively skipping irrelevant blocks at low overhead is still open. Existing data-skipping efforts maintain lightweight summaries (e.g., min/max, histograms) for each block to filter irrelevant data. However, such techniques ignore patterns in real-world data, enabling ineffective use of the storage budget and may cause serious false positives. This paper presents Sieve, a learning-enhanced index designed to efficiently filter out irrelevant blocks by capturing data patterns. Specifically, Sieve utilizes piece-wise linear functions to capture block distribution trends over the key space. Based on the captured trends, Sieve trades off storage consumption and false positives by grouping neighboring keys with similar block distributions into a single region. We have evaluated Sieve using Presto, and experiments on real-world datasets demonstrate that Sieve achieves up to 80% reduction in blocks accessed and 42% reduction in query times compared to its counterparts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Chameleon: Towards Update-Efficient Learned Indexing for Locally Skewed DataNa Guo, Yaqi Wang, Wenli Sun, Yu Gu 等ICDE 2024 · 被引用 6 次
- Optimizing Collections of Bloom Filters within a Space BudgetGabriel Mersy, Zhuo Wang, Stavros Sintos, Sanjay KrishnanVLDB 2024 · 被引用 2 次
- Robust Predicate Transfer with Dynamic ExecutionYiming Qiao, Peter Boncz, Huanchen ZhangVLDB 2026 · 被引用 2 次
它引用的顶会 Paper5
- ALEX: An Updatable Adaptive Learned IndexJialin Ding, Umar Farooq Minhas, Jia Yu, Chi Wang 等SIGMOD 2020 · 被引用 274 次
- The PGM-index: a fully-dynamic compressed learned index with provable worst-case boundsPaolo Ferragina, Giorgio VinciguerraVLDB 2020 · 被引用 178 次
- Pushing Data-Induced Predicates Through Joins in Big-Data ClustersLaurel J. Orr, Srikanth Kandula, Surajit ChaudhuriVLDB 2020 · 被引用 35 次
- The Price of Tailoring the Index to Your Data: Poisoning Attacks on Learned Index StructuresEvgenios M. Kornaropoulos, Silei Ren, Roberto TamassiaSIGMOD 2022 · 被引用 13 次
- Cuckoo Index: A Lightweight Secondary Index StructureAndreas Kipf, Damian Chromejko, Alexander Hall, Peter Boncz 等VLDB 2020 · 被引用 12 次
相关 Paper
- PTO: A Workload-driven Predictive Table Optimizer for Lakehouse SystemsVenkata Vamsikrishna Meduri, David Kreismann, Ronald Barber, Berthold ReinwaldSIGMOD 2026
- Pando: Enhanced Data Skipping with Logical Data PartitioningSivaprasad Sudhir, Wenbo Tao, Nikolay Pavlovich Laptev, Cyrille Habis 等VLDB 2023 · 被引用 14 次
- Instance-Optimized Data Layouts for Cloud Analytics WorkloadsJialin Ding, Umar Farooq Minhas, Badrish Chandramouli, Chi Wang 等SIGMOD 2021 · 被引用 37 次
- Towards Optimizing Storage Costs on the CloudKoyel Mukherjee, Raunak Shah, Shiv Kumar Saini, Karanpreet Singh 等ICDE 2023 · 被引用 8 次
- BtrBlocks: Efficient Columnar Compression for Data LakesMaximilian Kuschewski, David Sauerwein, Adnan Alhomssi, Viktor LeisSIGMOD 2023 · 被引用 47 次
