Optimizing Block Skipping for High-Dimensional Data with Learned Adaptive Curve
Xu Chen, Shuncheng Liu, Tong Yuan, Tao Ye, Kai Zeng, Han Su, Kai Zheng
Abstract
In the realm of big data and cloud analytics, efficiently managing and retrieving high-dimensional data presents a critical challenge. Traditional indexes often struggle with the storage overhead inherent in large datasets. There is a growing interest in the adoption of Small Materialize Aggregation (SMA) among cloud database vendors due to its ability to maintain lightweight block-level metadata, facilitating efficient block skipping. However, SMA performance relies heavily on data layout. This is especially critical in scenarios with wide tables containing hundreds of dimensions, where the curse of dimensionality exacerbates the issue. In this paper, we propose AdaCurve, a novel approach aimed at enhancing block skipping in high-dimensional datasets through adaptive optimization of data layout. Unlike conventional static and non-adaptive spacefilling curves (SFCs), AdaCurve leverages machine learning to develop an adaptive curve-a dynamically adjusting optimal projection function tailored to high-dimensional workloads and data characteristics. We introduce an attention-based network to handle high-dimensional data and a learnable objective for training adaptive curves in an end-to-end manner. Extensive experiments conducted on the Spark with real-world datasets demonstrate the effectiveness of AdaCurve. We have shown that AdaCurve effectively scales to datasets with dimensions of up to 1,000 columns, achieving a 2.8× improvement in block skipping compared to SFCs. CCS Concepts: • Information systems → Record and block layout.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64c5f745-90a8-43d2-8fcc-1c1ec18747aaBuilds on12
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Fast Differentiable Sorting and RankingMathieu Blondel, Olivier Teboul, Quentin Berthet, Josip DjolongaICML 2020 · 285 citations
- Deep Unsupervised Cardinality EstimationZongheng Yang, Eric Liang, Amog Kamsetty, Chenggang Wu et al.VLDB 2020 · 206 citations
- Learning Multi-Dimensional IndexesVikram Nathan, Jialin Ding, Mohammad Alizadeh, Tim KraskaSIGMOD 2020 · 180 citations
- Tsunami: A Learned Multi-dimensional Index for Correlated Data and Skewed WorkloadsJialin Ding, Vikram Nathan, Mohammad Alizadeh, Tim KraskaVLDB 2021 · 178 citations
Related papers
- LMSFC: A Novel Multidimensional Index based on Learned Monotonic Space Filling CurvesJian Gao, Xin Cao, Xin Yao, Gong Zhang et al.VLDB 2023 · 19 citations
- Sieve: A Learned Data-Skipping Index for Data AnalyticsYulai Tong, Jiazhen Liu, Hua Wang, Ke Zhou et al.VLDB 2023 · 10 citations
- Efficient Cost Modeling of Space-filling CurvesGuanli Liu, Lars Kulik, Christian S. Jensen, Tianyi Li et al.VLDB 2024 · 2 citations
- Instance-Optimized Data Layouts for Cloud Analytics WorkloadsJialin Ding, Umar Farooq Minhas, Badrish Chandramouli, Chi Wang et al.SIGMOD 2021 · 37 citations
- Pando: Enhanced Data Skipping with Logical Data PartitioningSivaprasad Sudhir, Wenbo Tao, Nikolay Pavlovich Laptev, Cyrille Habis et al.VLDB 2023 · 14 citations
