MemQ: A Graph-Based Query Memory Prediction Framework for Effective Workload Scheduling
Yang Wu, Xuanhe Zhou, Xiaoguang Li, Jinhuai Kang, Chunxiao Xing, Tongliang Li, Xinjun Yang, Wenchao Zhou, Feifei Li, Yong Zhang
Abstract
Query memory prediction is an essential yet underexplored problem in self-driving databases, particularly for high-concurrency workload scheduling where efficient resource utilization is critical. Existing works mainly focus on cost and latency estimation (e.g., using plan representation learning), while memory prediction poses new challenges such as requiring (1) numerous memory-specific training data, (2) memory-relevant query plan featurization strategies, and (3) a prediction model suitable for capturing the complexities of memory usage in query operations. Moreover, most learning-based approaches do not consider transferability across different datasets and database systems. This paper introduces Mem, a graph-based memory prediction framework designed for effective workload scheduling. First, we build a comprehensive training dataset for memory prediction by executing diverse query workloads across multiple datasets and recording their diverse peak memory consumptions. Second, our MemQ model leverages operator-level features of query plans, achieving high prediction accuracy, compact model size, and fast training and inference times. Third, we integrate the MemQ model into memory-aware First Fit Decreasing (FFD) and Bidrectional Fit (BF) scheduling strategy to optimize resource utilization. Extensive experiments demonstrate the effectiveness of our homogeneous query plan graph model. Moreover, our FFD scheduling strategy reduces makespan (total query execution time) by up to 55% and decreases retry counts by over 99% compared to default strategies when batch executing analytical queries on PostgreSQL. Furthermore, our novel BF strategy reduces makespan by 15.17% and reduces sum of total time by 41.41% compared with FFD strategy when batch executing mixed workloads.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers1
Ask how each one uses itRelated papers
- Query Performance Prediction for Concurrent Queries using Graph EmbeddingXuanhe Zhou, Ji Sun, Guoliang Li, Jianhua FengVLDB 2020 · 96 citations
- LSched: A Workload-Aware Learned Query Scheduler for Analytical Database SystemsIbrahim Sabek, Tenzin Samten Ukyab, Tim KraskaSIGMOD 2022 · 25 citations
- LAMP: A Dual-Mode Framework for Database Workload Memory PredictionGuoze Xue, Lu Chen, Ziquan Fang, Yushuai Li et al.ICDE 2026
- Tastes Great! Less Filling! High Performance and Accurate Training Data Collection for Self-Driving Database Management SystemsMatthew Butrovich, Wan Shen Lim, Lin Ma, John Rollinson et al.SIGMOD 2022 · 12 citations
- SafeLoad: Efficient Admission Control Framework for Identifying Memory-Overloading Queries in Cloud Data WarehousesYifan Wu, Yuhan Li, Zhenhua Wang, Zhongle Xie et al.VLDB 2026
