SEFRQO: A Self-Evolving Fine-Tuned RAG-Based Query Optimizer
Hanwen Liu, Qihan Zhang, Ryan Marcus, Ibrahim Sabek
Abstract
Query optimization is a crucial problem in database systems that has been studied for decades. Learned query optimizers (LQOs) can improve performance over time by incorporating feedback; however, they suffer from cold-start issues and often require retraining when workloads shift or schemas change. Recent LLM-based query optimizers leverage pre-trained and fine-tuned LLMs to mitigate these challenges. Nevertheless, they neglect LLMs' in-context learning and execution records as feedback for continuous evolution. In this paper, we present SEFRQO, a S elf- E volving F ine-tuned R AG-based Q uery O ptimizer. SEFRQO mitigates the cold-start problem of LQOs by continuously learning from execution feedback via a Retrieval-Augmented Generation (RAG) framework. We employ both supervised fine-tuning and reinforcement fine-tuning to prepare the LLM to produce syntactically correct and performance-efficient query hints. Moreover, SEFRQO leverages the LLM's in-context learning capabilities by dynamically constructing prompts with references to similar queries and the historical execution record of the same query. This self-evolving paradigm iteratively optimizes the prompt to minimize query execution latency. Evaluations show that SEFRQO outperforms state-of-the-art LQOs, achieving up to 65.05% and 93.57% reductions in query latency on the CEB and Stack workloads, respectively, compared to PostgreSQL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 19b50f30-9df2-4c45-a597-58a3e67601e6Cited by top-tier papers1
Ask how each one uses itBuilds on29
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Transformer in TransformerKai Han, An Xiao, Enhua Wu, Jianyuan Guo et al.NeurIPS 2021 · 2,148 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
- Text-to-SQL Empowered by Large Language Models: A Benchmark EvaluationDawei Gao, Haibin Wang, Yaliang Li, Xiuyu Sun et al.VLDB 2024 · 609 citations
Related papers
- LLM4Hint: Leveraging Large Language Models for Hint Recommendation in Offline Query OptimizationSuchen Liu, Yang Lin, Yinjun Han, Jun GaoICDE 2026 · 1 citation
- Can Large Language Models Be Query Optimizer for Relational Databases?Jie Tan, Kangfei Zhao, Rui Li, Jeffrey Xu Yu et al.SIGMOD 2026 · 6 citations
- LIMAO: A Framework for Lifelong Modular Learned Query OptimizationQihan Zhang, Shaolin Xie, Ibrahim SabekVLDB 2025 · 4 citations
- Divo: Learning a Stable and Effective Query Optimizer with a Diverse WorkloadTianyi Chen, Jun Gao, Yaofeng Tu, Yang Lin et al.SIGMOD 2026 · 1 citation
- Lequa: A Learning-Based Query-Aware Framework for Selective Query OptimizationGuoneng Li, Pengfei Zheng, Ling Xu, Yan Li et al.ICDE 2026
