SageDB: An Instance-Optimized Data Analytics System
Jialin Ding, Ryan Marcus, Andreas Kipf, Vikram Nathan, Aniruddha Nrusimha, Kapil Vaidya, Alexander van Renen, Tim Kraska
Abstract
Modern data systems are typically both complex and general-purpose. They are complex because of the numerous internal knobs and parameters that users need to manually tune in order to achieve good performance; they are general-purpose because they are designed to handle diverse use cases, and therefore often do not achieve the best possible performance for any specific use case. A recent trend aims to tackle these pitfalls: instance-optimized systems are designed to automatically self-adjust in order to achieve the best performance for a specific use case, i.e., a dataset and query workload. Thus far, the research community has focused on creating instance-optimized database components, such as learned indexes and learned cardinality estimators, which are evaluated in isolation. However, to the best of our knowledge, there is no complete data system built with instance-optimization as a foundational design principle. In this paper, we present a progress report on SageDB, our effort towards building the first instance-optimized data system. SageDB synthesizes various instance-optimization techniques to automatically specialize for a given use case, while simultaneously exposing a simple user interface that places minimal technical burden on the user. Our prototype outperforms a commercial cloud-based analytics system by up to 3× on end-to-end query workloads and up to 250× on individual queries. SageDB is an ongoing research effort, and we highlight our lessons learned and key directions for future work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext be8181eb-50ab-4005-9047-6bf7ac94db9cCited by top-tier papers2
- Grep: A Graph Learning Based Database Partitioning SystemXuanhe Zhou, Guoliang Li, Jianhua Feng, Luyang Liu et al.SIGMOD 2023 · 14 citations
- A New Paradigm in Tuning Learned Indexes: A Reinforcement Learning Enhanced ApproachTaiyi Wang, Liang Liang, Guang Yang, Thomas Heinis et al.SIGMOD 2025 · 2 citations
Builds on10
- ALEX: An Updatable Adaptive Learned IndexJialin Ding, Umar Farooq Minhas, Jia Yu, Chi Wang et al.SIGMOD 2020 · 274 citations
- Bao: Making Learned Query Optimization PracticalRyan Marcus, Parimarjan Negi, Hongzi Mao, Nesime Tatbul et al.SIGMOD 2021 · 242 citations
- Learning Multi-Dimensional IndexesVikram Nathan, Jialin Ding, Mohammad Alizadeh, Tim KraskaSIGMOD 2020 · 180 citations
- Tsunami: A Learned Multi-dimensional Index for Correlated Data and Skewed WorkloadsJialin Ding, Vikram Nathan, Mohammad Alizadeh, Tim KraskaVLDB 2021 · 178 citations
- From WiscKey to Bourbon: A Learned Index for Log-Structured Merge TreesYifan Dai, Yien Xu, Aishwarya Ganesan, Ramnatthan Alagappan et al.OSDI 2020 · 138 citations
Related papers
- Data-Agnostic Cardinality Learning from Imperfect WorkloadsPeizhi Wu, Rong Kang, Tieying Zhang, Jianjun Chen et al.VLDB 2025 · 1 citation
- Check Out the Big Brain on BRAD: Simplifying Cloud Data Processing with Learned Automated Data MeshesTim Kraska, Tianyu Li, Samuel Madden, Markos Markakis et al.VLDB 2023 · 13 citations
- Practical Parameterized Query Optimization via Efficient Plan Reuse and List-wise RankingHai Lan, Yang Yu, Zhifeng Bao, Zi Huang et al.SIGMOD 2026
- LSched: A Workload-Aware Learned Query Scheduler for Analytical Database SystemsIbrahim Sabek, Tenzin Samten Ukyab, Tim KraskaSIGMOD 2022 · 25 citations
- Towards Cost-Optimal Query Processing in the CloudViktor Leis, Maximilian KuschewskiVLDB 2021 · 34 citations
