Replicated Layout for In-Memory Database Systems
Sivaprasad Sudhir, Michael J. Cafarella, Samuel Madden
Abstract
Scanning and filtering are the foundations of analytical database systems. Modern DBMSs employ a variety of techniques to partition and layout data to improve the performance of these operations. To accelerate query performance, systems tune data layout to reduce the cost of accessing and processing data. However, these layouts optimize for the average query, and with heterogeneous data access patterns in parts of the data, their performance degrades. To mitigate this, we present CopyRight, a layout-aware partial replication engine that replicates parts of the data differently and lays out each replica differently to maximize the overall query performance. Across a range of real-world query workloads, CopyRight is able to achieve 1.1X to 7.9X faster performance than the best non-replicated layout with 0.25X space overhead. When compared to full table replication with 100% overhead, CopyRight attains the same or up to 5.2X speedup with 25% space overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c97f12b3-cb34-47ca-8a8a-3acd3c43ad4eCited by top-tier papers4
- SageDB: An Instance-Optimized Data Analytics SystemJialin Ding, Ryan Marcus, Andreas Kipf, Vikram Nathan et al.VLDB 2022 · 18 citations
- Pando: Enhanced Data Skipping with Logical Data PartitioningSivaprasad Sudhir, Wenbo Tao, Nikolay Pavlovich Laptev, Cyrille Habis et al.VLDB 2023 · 14 citations
- Blitzcrank: Fast Semantic Compression for In-memory Online Transaction ProcessingYiming Qiao, Yihan Gao, Huanchen ZhangVLDB 2024 · 3 citations
- Benchmarking Adaptive Multidimensional IndicesKonstantinos Lampropoulos, Fatemeh Zardbani, Nikos Mamoulis, Panagiotis KarrasVLDB 2025
Builds on7
- Learning Multi-Dimensional IndexesVikram Nathan, Jialin Ding, Mohammad Alizadeh, Tim KraskaSIGMOD 2020 · 180 citations
- Tsunami: A Learned Multi-dimensional Index for Correlated Data and Skewed WorkloadsJialin Ding, Vikram Nathan, Mohammad Alizadeh, Tim KraskaVLDB 2021 · 178 citations
- LISA: A Learned Index Structure for Spatial DataPengfei Li, Hua Lu, Qian Zheng, Long Yang et al.SIGMOD 2020 · 158 citations
- Effectively Learning Spatial IndicesJianzhong Qi, Guanli Liu, Christian S. Jensen, Lars KulikVLDB 2020 · 121 citations
- Qd-tree: Learning Data Layouts for Big Data AnalyticsZongheng Yang, Badrish Chandramouli, Chi Wang, Johannes Gehrke et al.SIGMOD 2020 · 87 citations
Related papers
- Scaling your Hybrid CPU-GPU DBMS to Multiple GPUsBobbi W. Yogatama, Weiwei Gong, Xiangyao YuVLDB 2024 · 9 citations
- Dynamic Data Layout Optimization with Worst-Case GuaranteesKexin Rong, Paul Liu, Sarah Ashok Sonje, Moses CharikarICDE 2024 · 1 citation
- Rethink Query Optimization in HTAP DatabasesHaoze Song, Wenchao Zhou, Feifei Li, Xiang Peng et al.SIGMOD 2024 · 7 citations
- Proteus: Autonomous Adaptive Storage for Mixed WorkloadsMichael Abebe, Horatiu Lazu, Khuzaima DaudjeeSIGMOD 2022 · 20 citations
- Differentially Oblivious Relational Database OperatorsLianke Qin, Rajesh Jayaram, Elaine Shi, Zhao Song et al.VLDB 2023 · 12 citations
