TATA: An Efficient Framework for Task Transfer in Query Plan Representation
Yue Zhao, Songsong Mo, Gao Cong
Abstract
Machine learning for database systems has achieved significant success in various database components, such as cost estimation, query optimization, index selection, view recommendation, and semantic equivalence detection. However, these solutions typically focus on a single task and normally need a large amount of labeled data for the task to train machine learning models. Even if a solution can be adapted for a different task, it will require recollecting labeled data for each new task, which is typically much more time-consuming than model training. While dataset collection is relatively easier for some tasks, it can be prohibitively expensive for others. A natural solution is to use transfer learning techniques to adapt learned knowledge from one task to another. However, we show that naive transfer learning methods perform poorly and are only as good as training from scratch. Their failures are mainly due to three challenges: (1) the source model is not robust as it is optimized to its task only; (2) the size of the target dataset is small; and (3) the inevitable distribution shift when changing tasks. To overcome these challenges, we first study the task transfer problem in query plan representation and propose a new framework TATA for the problem. Specifically, to address the lack of robustness in the source model, TATA incorporates a self-supervised component during the pretraining stage. Specifically, we design a query plan decoder to reconstruct the original query plan from its representation, ensuring the model preserves key features. This leads to more robust and transferable query plan representations. Next, to address the issues of small datasets and distribution shift, TATA generates an arbitrary number of query plans for the target task and assigns them realistic pseudo labels. This is achieved by utilizing both strong database domain knowledge and available datasets. Through extensive experiments, we show that TATA delivers substantial improvements on task transfer, achieving up to 5× reduction in dataset collection cost when transferring from cost estimation to two representative target tasks: query optimization and index selection. We demonstrate compatibility with three distinct query plan representation models, establishing broader applicability than prior transfer approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f117a2be-686c-4f91-8f20-c2fa9cb227bdBuilds on24
- An End-to-End Learning-based Cost EstimatorJi Sun, Guoliang LiVLDB 2020 · 251 citations
- Bao: Making Learned Query Optimization PracticalRyan Marcus, Parimarjan Negi, Hongzi Mao, Nesime Tatbul et al.SIGMOD 2021 · 242 citations
- Reinforcement Learning with Tree-LSTM for Join Order SelectionXiang Yu, Guoliang Li, Chengliang Chai, Nan TangICDE 2020 · 168 citations
- QueryFormer: A Tree Transformer Model for Query Plan RepresentationYue Zhao, Gao Cong, Jiachen Shi, Chunyan MiaoVLDB 2022 · 117 citations
- Lero: A Learning-to-Rank Query OptimizerRong Zhu, Wei Chen, Bolin Ding, Xingguang Chen et al.VLDB 2023 · 102 citations
Related papers
- A Comparative Study and Component Analysis of Query Plan Representation Techniques in ML4DB StudiesYue Zhao, Zhaodonghui Li, Gao CongVLDB 2024 · 19 citations
- PRICE: A Pretrained Model for Cross-Database Cardinality EstimationTianjing Zeng, Junwei Lan, Jiahong Ma, Wenqing Wei et al.VLDB 2025 · 14 citations
- Robust Plan Evaluation based on Approximate Probabilistic Machine LearningAmin Kamali, Verena Kantere, Calisto Zuzarte, Vincent CorvinelliVLDB 2025 · 1 citation
- Expand your Training Limits! Generating Training Data for ML-based Data ManagementFrancesco Ventura, Zoi Kaoudi, Jorge-Arnulfo Quiané-Ruiz, Volker MarklSIGMOD 2021 · 19 citations
- Detect, Distill and Update: Learned DB Systems Facing Out of Distribution DataMeghdad Kurmanji, Peter TriantafillouSIGMOD 2023 · 19 citations
