Parallel kd-tree with Batch Updates
Ziyang Men, Zheqi Shen, Yan Gu, Yihan Sun
Abstract
The 𝑘d-tree is one of the most widely used data structures to manage multi-dimensional data. Due to the ever-growing data volume, it is imperative to consider parallelism in 𝑘d-trees. However, we observed challenges in existing parallel 𝑘d-tree implementations, for both constructions and updates.
The goal of this paper is to develop efficient in-memory 𝑘d-trees by supporting high parallelism and cache-efficiency. We propose the Pkd-tree (Parallel 𝑘d-tree), a parallel 𝑘d-tree that is efficient both in theory and in practice. The Pkd-tree supports parallel tree construction, batch update (insertion and deletion), and various queries including 𝑘-nearest neighbor search, range query, and range count. We proved that our algorithms have strong theoretical bounds in work (sequential time complexity), span (parallelism), and cache complexity. Our key techniques include 1) an efficient construction algorithm that optimizes work, span, and cache complexity simultaneously, and 2) reconstruction-based update algorithms that guarantee the tree to be weight-balanced. With the new algorithmic insights and careful engineering effort, we achieved a highly optimized implementation of the Pkd-tree.
We tested Pkd-tree with various synthetic and real-world datasets, including both uniform and highly skewed data. We compare the Pkd-tree with state-of-the-art parallel 𝑘d-tree implementations. In all tests, with better or competitive query performance, Pkd-tree is much faster in construction and updates consistently than all baselines. We released our code.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 19614f33-dc4e-450d-918b-9e5a9a9982d7Cited by top-tier papers3
- PIM-zd-tree: A Fast Space-Partitioning Index Leveraging Processing-in-MemoryYiwei Zhao, Hongbo Kang, Ziyang Men, Yan Gu et al.PPoPP 2026 · 1 citation
- Parallel Dynamic Spatial IndexesZiyang Men, Bo Huang, Yan Gu, Yihan SunPPoPP 2026 · 1 citation
- Trading Vector Data in Vector DatabasesJin Cheng, Xiangxiang Dai, Ningning Ding, John C. S. Lui et al.ICDE 2026
Builds on4
- Fast Parallel Algorithms for Euclidean Minimum Spanning Tree and Hierarchical Spatial ClusteringYiqiu Wang, Shangdi Yu, Yan Gu, Julian ShunSIGMOD 2021 · 34 citations
- PaC-trees: supporting parallel and compressed purely-functional collectionsLaxman Dhulipala, Guy E. Blelloch, Yan Gu, Yihan SunPLDI 2022 · 16 citations
- Parallel Integer Sort: Theory and PracticeXiaojun Dong, Laxman Dhulipala, Yan Gu, Yihan SunPPoPP 2024 · 9 citations
- A Scalable and Generic Approach to Range JoinsMaximilian Reif, Thomas NeumannVLDB 2022 · 6 citations
Related papers
- UFO Trees: Practical and Provably-Efficient Parallel Batch-Dynamic TreesQuinten De Man, Atharva Sharma, Kishen N. Gowda, Laxman DhulipalaPPoPP 2026 · 1 citation
- Tree Embedding in High Dimensions: Dynamic and Massively ParallelGramoz Goranci, Shaofeng H.-C. Jiang, Peter Kiss, Qihao Kong et al.SODA 2026
- ParallelNN: A Parallel Octree-based Nearest Neighbor Search Accelerator for 3D Point CloudsFaquan Chen, Rendong Ying, Jianwei Xue, Fei Wen et al.HPCA 2023 · 37 citations
- PIM-tree: A Skew-resistant Index for Processing-in-MemoryHongbo Kang, Yiwei Zhao, Guy E. Blelloch, Laxman Dhulipala et al.VLDB 2023 · 40 citations
- Evaluating Persistent Memory Range IndexesLucas Lersch, Xiangpeng Hao, Ismail Oukid, Tianzheng Wang et al.VLDB 2020 · 97 citations
