Columnar Storage and List-based Processing for Graph Database Management Systems
Pranjal Gupta, Amine Mhedhbi, Semih Salihoglu
Abstract
We revisit column-oriented storage and query processing techniques in the context of contemporary graph database management systems (GDBMSs). Similar to column-oriented RDBMSs, GDBMSs support read-heavy analytical workloads that however have fundamentally different data access patterns than traditional analytical workloads. We first derive a set of desiderata for optimizing storage and query processors of GDBMS based on their access patterns. We then present the design of columnar storage, compression, and query processing techniques based on these desiderata. In addition to showing direct integration of existing techniques from columnar RDBMSs, we also propose novel ones that are optimized for GDBMSs. These include a novel list-based query processor, which avoids expensive data copies of traditional block-based processors under many-to-many joins, a new data structure we call single-indexed edge property pages and an accompanying edge ID scheme, and a new application of Jacobson's bit vector index for compressing NULL values and empty lists. We integrated our techniques into the GraphflowDB in-memory GDBMS. Through extensive experiments, we demonstrate the scalability and query performance benefits of our techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 601942fd-54d9-4aa6-9c25-c87306eab24aCited by top-tier papers12
- The LDBC Social Network Benchmark: Business Intelligence WorkloadGábor Szárnyas, Jack Waudby, Benjamin A. Steer, Dávid Szakállas et al.VLDB 2023 · 103 citations
- NaviX: A Native Vector Index Design for Graph DBMSs With Robust Predicate-Agnostic Search PerformanceGaurav Sehgal, Semih SalihogluVLDB 2025 · 12 citations
- Grouping Time Series for Efficient Columnar StorageChenguang Fang, Shaoxu Song, Haoquan Guan, Xiangdong Huang et al.SIGMOD 2023 · 10 citations
- AeonG: An Efficient Built-in Temporal Support in Graph DatabasesJiamin Hou, Zhanhao Zhao, Zhouyu Wang, Wei Lu et al.VLDB 2024 · 8 citations
- GraphAr: An Efficient Storage Scheme for Graph Data in Data LakesXue Li, Weibin Zeng, Zhibin Wang, Diwen Zhu et al.VLDB 2025 · 5 citations
Builds on1
Related papers
- Making RDBMSs Efficient on Graph Workloads Through Predefined JoinsGuodong Jin, Semih SalihogluVLDB 2022 · 24 citations
- A+ Indexes: Tunable and Space-Efficient Adjacency Lists in Graph Database Management SystemsAmine Mhedhbi, Pranjal Gupta, Shahid Khaliq, Semih SalihogluICDE 2021 · 10 citations
- MorphStore: Analytical Query Engine with a Holistic Compression-Enabled Processing ModelPatrick Damme, Annett Ungethüm, Johannes Pietrzyk, Alexander Krause et al.VLDB 2020
- Enabling Index-free Adjacency in Oblivious Graph Processing with Delayed DuplicationsWeiqi Feng, Xinle Cao, Adam O'Neill, Chuanhui YangVLDB 2026
- Optimizing Differentially-Maintained Recursive Queries on Dynamic GraphsKhaled Ammar, Siddhartha Sahu, Semih Salihoglu, M. Tamer ÖzsuVLDB 2022 · 6 citations
