Mainlining Databases: Supporting Fast Transactional Workloads on Universal Columnar Data File Formats
Tianyu Li, Matthew Butrovich, Amadou Ngom, Wan Shen Lim, Wes McKinney, Andrew Pavlo
Abstract
The proliferation of modern data processing tools has given rise to open-source columnar data formats. These formats help organizations avoid repeated conversion of data to a new format for each application. However, these formats are read-only, and organizations must use a heavy-weight transformation process to load data from on-line transactional processing (OLTP) systems. As a result, DBMSs often fail to take advantage of full network bandwidth when transferring data. We aim to reduce or even eliminate this overhead by developing a storage architecture for in-memory database management systems (DBMSs) that is aware of the eventual usage of its data and emits columnar storage blocks in a universal open-source format. We introduce relaxations to common analytical data formats to efficiently update records and rely on a lightweight transformation process to convert blocks to a read-optimized layout when they are cold. We also describe how to access data from third-party analytical tools with minimal serialization overhead. We implemented our storage engine based on the Apache Arrow format and integrated it into the NoisePage DBMS to evaluate our work. Our experiments show that our approach achieves comparable performance with dedicated OLTP DBMSs while enabling orders-of-magnitude faster data exports to external data science and machine learning tools than existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0308950a-3577-46fb-8a42-cf206d98dce4Cited by top-tier papers8
- Permutable Compiled Queries: Dynamically Adapting Compiled Queries without RecompilingPrashanth Menon, Amadou Ngom, Todd C. Mowry, Andrew Pavlo et al.VLDB 2021 · 26 citations
- A Deep Dive into Common Open Formats for Analytical DBMSsChunwei Liu, Anna Pavlenko, Matteo Interlandi, Brandon HaynesVLDB 2023 · 23 citations
- ConnectorX: Accelerating Data Loading From Databases to DataframesXiaoying Wang, Weiyuan Wu, Jinze Wu, Yizhou Chen et al.VLDB 2022 · 14 citations
- Deploying Computational Storage for HTAP DBMSs Takes More Than Just Computation OffloadingKitaek Lee, Insoon Jo, Jaechan Ahn, Hyuk Lee et al.VLDB 2023 · 14 citations
- Scalable and Robust Snapshot Isolation for High-Performance Storage EnginesAdnan Alhomssi, Viktor LeisVLDB 2023 · 12 citations
Related papers
- An Empirical Evaluation of Columnar Storage FormatsXinyu Zeng, Yulong Hui, Jiahong Shen, Andrew Pavlo et al.VLDB 2024 · 59 citations
- MorphStore: Analytical Query Engine with a Holistic Compression-Enabled Processing ModelPatrick Damme, Annett Ungethüm, Johannes Pietrzyk, Alexander Krause et al.VLDB 2020
- These Rows Are Made for Sorting and That's Just What We'll DoLaurens Kuiper, Hannes MühleisenICDE 2023 · 6 citations
- Columnar Storage and List-based Processing for Graph Database Management SystemsPranjal Gupta, Amine Mhedhbi, Semih SalihogluVLDB 2021 · 31 citations
- BOSS - An Architecture for Database Kernel CompositionHubert Mohr-Daurat, Xuan Sun, Holger PirkVLDB 2024 · 12 citations
