LATEST: Learning-Assisted Selectivity Estimation Over Spatio-Textual Streams
Mayur Patil, Amr Magdy
Abstract
Selectivity and cardinality estimation are main driving factors for developing cheap query plans and ultimately faster query processing. Traditionally, database systems use estimation data structures, e.g., histograms, to maintain data summaries. Machine learning models have recently been employed, acting as black boxes, in several database tasks, including cardinality estimation. In the dynamic streaming environments, both estimation data structures and machine learning models struggle with adaptation for dynamic changes in data and query workloads. This paper proposes LATEST; a system module that uses machine learning to enable dynamic adaptation of estimation data structures. For spatial-keyword queries in a streaming environment, it shows on par or better performance than the state-of-the-art estimators. LATEST builds an incremental supervised learning model over a moving time window that helps the underlying system to switch among several estimation structures to keep estimation accuracy high at all times. As an incremental learner, LATEST effectively adapts to dynamic changes of both data and queries in streaming environments. Our extensive experiments on three real datasets and various query workloads verify the effectiveness of LATEST with higher accuracy and lower response times over the state-of-the-art estimators.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aafa0b94-1e8b-4f1b-a829-12fdc55fee5eBuilds on1
Related papers
- One Seed, Two Birds: A Unified Learned Structure for Exact and Approximate CountingYingze Li, Hongzhi Wang, Xianglong LiuSIGMOD 2024 · 4 citations
- Warper: Efficiently Adapting Learned Cardinality Estimators to Data and Workload DriftsBeibin Li, Yao Lu, Srikanth KandulaSIGMOD 2022 · 29 citations
- DeepDB: Learn from Data, not from Queries!Benjamin Hilprecht, Andreas Schmidt, Moritz Kulessa, Alejandro Molina et al.VLDB 2020 · 154 citations
- LHist: Towards Learning Multi-dimensional Histogram for Massive Spatial DataQiyu Liu, Yanyan Shen, Lei ChenICDE 2021 · 20 citations
- QuickSel: Quick Selectivity Learning with Mixture ModelsYongjoo Park, Shucheng Zhong, Barzan MozafariSIGMOD 2020 · 66 citations
