PIM-tree: A Skew-resistant Index for Processing-in-Memory
Hongbo Kang, Yiwei Zhao, Guy E. Blelloch, Laxman Dhulipala, Yan Gu, Charles McGuffey, Phillip B. Gibbons
Abstract
The performance of today’s in-memory indexes is bottlenecked by the memory latency/bandwidth wall. Processing-in-memory (PIM) is an emerging approach that potentially mitigates this bottleneck by enabling low-latency memory access whose aggregate memory bandwidth scales with the number of PIM nodes. There is an inherent tension, however, between minimizing inter-node communication and achieving load balance in PIM systems, in the presence of workload skew. This paper presents PIM-tree, an ordered index for PIM systems that achieves both low communication and high load balance, regardless of the degree of skew in data/queries. Our skew-resistant index is based on a novel division of labor between the multi-core host CPU and the PIM nodes, which leverages the strengths of each. We introduce push-pull search, which dynamically decides whether to push queries to a PIM-tree node (CPU →minimal amsmath wasysym amsfonts amssymb amsbsy mathrsfs upgreek -69pt documentdocument PIM-node) or pull the node’s keys back to the CPU (PIM-node →minimal amsmath wasysym amsfonts amssymb amsbsy mathrsfs upgreek -69pt documentdocument CPU) based on workload skew. Combined with other PIM-friendly optimizations (shadow subtrees and chunking), PIM-tree achieves high throughput, (guaranteed) low communication, and (guaranteed) high load balance, for batches of point queries, updates, and range scans. We implement the PIM-tree structure, in addition to prior proposed PIM indexes, on the latest PIM system from UPMEM, with 32 CPU cores and 2048 PIM nodes. On workloads with 500 million keys and batches of 1 million queries, the throughput using PIM-trees is up to 69.7×minimal amsmath wasysym amsfonts amssymb amsbsy mathrsfs upgreek -69pt documentdocument and 59.1×minimal amsmath wasysym amsfonts amssymb amsbsy mathrsfs upgreek -69pt documentdocument higher than the two best prior PIM-based methods. As far as we know these are the first implementations of ordered indexes on real PIM systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9a17f69-8e1b-4e23-b9ad-62ec0fb0a87fCited by top-tier papers15
- PIM-DL: Expanding the Applicability of Commodity DRAM-PIMs for Deep Learning via Algorithm-System Co-OptimizationCong Li, Zhe Zhou, Yang Wang, Fan Yang et al.ASPLOS 2024 · 27 citations
- NDPBridge: Enabling Cross-Bank Coordination in Near-DRAM-Bank Processing ArchitecturesBoyu Tian, Yiwei Li, Li Jiang, Shuangyu Cai et al.ISCA 2024 · 27 citations
- PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM SystemsDongjae Lee, Bongjoon Hyun, Taehun Kim, Minsoo RhuMICRO 2024 · 23 citations
- PimPam: Efficient Graph Pattern Matching on Real Processing-in-Memory HardwareShuangyu Cai, Boyu Tian, Huanchen Zhang, Mingyu GaoSIGMOD 2024 · 18 citations
- Bf-Tree: A Modern Read-Write-Optimized Concurrent Larger-Than-Memory Range IndexXiangpeng Hao, Badrish ChandramouliVLDB 2024 · 14 citations
Builds on5
- SpaceA: Sparse Matrix Vector Multiplication on Processing-in-Memory AcceleratorXinfeng Xie, Zheng Liang, Peng Gu, Abanti Basak et al.HPCA 2021 · 111 citations
- SynCron: Efficient Synchronization Support for Near-Data-Processing ArchitecturesChristina Giannoula, Nandita Vijaykumar, Nikela Papadopoulou, Vasileios Karakostas et al.HPCA 2021 · 70 citations
- On Supporting Efficient Snapshot Isolation for Hybrid Workloads with Multi-Versioned IndexesYihan Sun, Guy E. Blelloch, Wan Shen Lim, Andrew PavloVLDB 2020 · 37 citations
- PIM-Assembler: A Processing-in-Memory Platform for Genome AssemblyShaahin Angizi, Naima Ahmed Fahmi, Wei Zhang, Deliang FanDAC 2020 · 27 citations
- PIM-Quantifier: A Processing-in-Memory Platform for mRNA QuantificationFan Zhang, Shaahin Angizi, Naima Ahmed Fahmi, Wei Zhang et al.DAC 2021 · 18 citations
Related papers
- PIM-zd-tree: A Fast Space-Partitioning Index Leveraging Processing-in-MemoryYiwei Zhao, Hongbo Kang, Ziyang Men, Yan Gu et al.PPoPP 2026 · 1 citation
- PIMLex: A High-Performance Learned Index with Processing-in-MemoryLixiao Cui, Kedi Yang, Yusen Li, Gang Wang et al.FAST 2025 · 11 citations
- A Case Study of Processing-in-Memory in off-the-Shelf SystemsJoel Nider, Craig Mustard, Andrada Zoltan, John Ramsden et al.USENIX ATC 2021 · 62 citations
- PIMnet: A Domain-Specific Network for Efficient Collective Communication in Scalable PIMHyojun Son, Gilbert Jonatan, Xiangyu Wu, Haeyoon Cho et al.HPCA 2025 · 7 citations
- Turbocharge ANNS on Real Processing-in-Memory by Enabling Fine-Grained Per-PIM-Core SchedulingPuqing Wu, Minhui Xie, Enrui Zhao, Dafang Zhang et al.USENIX ATC 2025 · 8 citations
