Towards Systematic Index Dynamization
Douglas B. Rumbaugh, Dong Xie, Zhuoyue Zhao
Abstract
There is significant interest in examining large datasets using complex domain-specific queries. In many cases, these queries can be accelerated using specialized indexes. Unfortunately, the development of a practical index is difficult, because databases generally require additional features such as updates, concurrency support, crash recovery, etc. There are three major lines of work to alleviate the pain: (1) automatic index composition/tuning which composes indexes out of core data structure primitives to optimize for specific workloads; (2) generalized index templates which generalize common data structures such as B+-trees for custom queries over custom data types, and (3) data structure dynamization frameworks such as the Bentley-Saxe method which converts a static data structure into an updatable data structure with bounded additional query cost. The first two are limited to very specific queries and/or data structures and, thus, are not suitable for building a general index dynamization framework. The last one is more promising in its generality but also has limitations on query types, deletion support, and performance tuning. In this paper, we discuss the limitations of the classic index dynamization techniques and propose a path towards a more general and systematic solution. We demonstrate the viability of our framework by realizing it as a C++20 metaprogramming library and conducting case studies on four example queries with their corresponding static index structures. With this framework, many theoretical/early-stage index designs can easily be extended with support for updates, along with a wide tuning space for query/update performance trade-offs. This allows index designers to focus on efficient data layouts and query algorithms, thereby dramatically narrowing the gap between novel index designs and deployment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 05a06ccd-e91e-4b0d-a1c6-3c6b04ad4831Cited by top-tier papers1
Ask how each one uses itBuilds on9
- ALEX: An Updatable Adaptive Learned IndexJialin Ding, Umar Farooq Minhas, Jia Yu, Chi Wang et al.SIGMOD 2020 · 274 citations
- The PGM-index: a fully-dynamic compressed learned index with provable worst-case boundsPaolo Ferragina, Giorgio VinciguerraVLDB 2020 · 178 citations
- Cosine: A Cloud-Cost Optimized Self-Designing Key-Value Storage EngineSubarna Chatterjee, Meena Jagadeesan, Wilson Qin, Stratos IdreosVLDB 2022 · 17 citations
- Spatial Independent Range SamplingDong Xie, Jeff M. Phillips, Michael Matheny, Feifei LiSIGMOD 2021 · 13 citations
- The next 50 Years in Database Indexing or: The Case for Automatically Generated Index StructuresJens Dittrich, Joris Nix, Christian SchönVLDB 2022 · 12 citations
Related papers
- On Self-Designing Learned IndexesBaofu Han, Guoyu Hu, Bing Li, Xiaokui Xiao et al.SIGMOD 2026
- Charting the Design Space of Query Execution using VOILATim Gubner, Peter BonczVLDB 2021 · 17 citations
- Adaptive Code Generation for Data-Intensive AnalyticsWangda Zhang, Junyoung Kim, Kenneth A. Ross, Eric Sedlar et al.VLDB 2021 · 12 citations
- Practical Dynamic Extension for Sampling IndexesDouglas B. Rumbaugh, Dong XieSIGMOD 2024 · 3 citations
- Query Compilation Without RegretsPhilipp M. Grulich, Aljoscha P. Lepping, Dwi P. A. Nugroho, Varun Pandey et al.SIGMOD 2024 · 6 citations
