Comprehensive and Efficient Workload Compression
Shaleen Deep, Anja Gruenheid, Paraschos Koutris, Jeffrey F. Naughton, Stratis Viglas
Abstract
This work studies the problem of constructing a representative workload from a given input analytical query workload where the former serves as an approximation with guarantees of the latter. We discuss our work in the context of workload analysis and monitoring. As an example, evolving system usage patterns in a database system can cause load imbalance and performance regressions which can be controlled by monitoring system usage patterns, i.e., a representative workload, over time. To construct such a workload in a principled manner, we formalize the notions of workload representativity and coverage. These metrics capture the intuition that the distribution of features in a compressed workload should match a target distribution, increasing representativity, and include common queries as well as outliers, increasing coverage. We show that solving this problem optimally is computationally hard and present a novel greedy algorithm that provides approximation guarantees. We compare our techniques to established algorithms in this problem space such as sampling and clustering, and demonstrate advantages and key trade-offs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7a6b693b-95a4-40ba-9a70-cf979c5fe411Cited by top-tier papers9
- DISTILL: Low-Overhead Data-Driven Techniques for Filtering and Costing Indexes for Scalable Index TuningTarique Siddiqui, Wentao Wu, Vivek R. Narasayya, Surajit ChaudhuriVLDB 2022 · 36 citations
- Budget-aware Index Tuning with Reinforcement LearningWentao Wu, Chi Wang, Tarique Siddiqui, Junxiong Wang et al.SIGMOD 2022 · 33 citations
- ISUM: Efficiently Compressing Large and Complex Workloads for Scalable Index TuningTarique Siddiqui, Saehan Jo, Wentao Wu, Chi Wang et al.SIGMOD 2022 · 25 citations
- Refactoring Index Tuning Process with Benefit EstimationTao Yu, Zhaonian Zou, Weihua Sun, Yu YanVLDB 2024 · 13 citations
- "What makes my queries slow?": Subgroup Discovery for SQL Workload AnalysisYoucef Remil, Anes Bendimerad, Romain Mathonat, Philippe Chaleat et al.ASE 2021 · 12 citations
Related papers
- DBAugur: An Adversarial-based Trend Forecasting System for Diversified WorkloadsYuanning Gao, Xiuqi Huang, Xuanhe Zhou, Xiaofeng Gao et al.ICDE 2023 · 10 citations
- Weighted Set Multi-Cover on Bounded Universe and Applications in Package RecommendationNima Shahbazi, Aryan Esmailpour, Stavros SintosSIGMOD 2026
- Modeling Shifting Workloads for Learned Database SystemsPeizhi Wu, Zachary G. IvesSIGMOD 2024 · 12 citations
- Understanding Queries by Conditional InstancesAmir Gilad, Zhengjie Miao, Sudeepa Roy, Jun YangSIGMOD 2022 · 9 citations
- Are Learned DBMS Components Robust to Workload Drift?: [Experiments & Analysis]Zizhong Meng, Gao Cong, Siqiang LuoSIGMOD 2026
