PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM Devices
Si Ung Noh, Junguk Hong, Chaemin Lim, Seongyeon Park, Jeehyun Kim, Hanjun Kim, Youngsok Kim, Jinho Lee
Abstract
Recent dual in-line memory modules (DIMMs) are starting to support processing-in-memory (PIM) by associating their memory banks with processing elements (PEs), allowing applications to overcome the data movement bottleneck by offloading memory-intensive operations to the PEs. Many highly parallel applications have been shown to benefit from these PIM-enabled DIMMs, but further speedup is often limited by the huge overhead of inter-PE collective communication. This mainly comes from the slow CPU-mediated inter-PE communication methods, making it difficult for PIM-enabled DIMMs to accelerate a wider range of applications. Prior studies have tried to alleviate the communication bottleneck, but they lack enough flexibility and performance to be used for a wide range of applications. In this paper, we present PID-Comm, a fast and flexible inter-PE collective communication framework for commodity PIM-enabled DIMMs. The key idea of PID-Comm is to abstract the PEs as a multi-dimensional hypercube and allow multiple instances of inter-PE collective communication between the PEs belonging to certain dimensions of the hypercube. Leveraging this abstraction, PID-Comm first defines eight interPE collective communication patterns that allow applications to easily express their complex communication patterns. Then, PIDComm provides high-performance implementations of the interPE collective communication patterns optimized for the DIMMs. Our evaluation using 16 UPMEM DIMMs and representative parallel algorithms shows that PID-Comm greatly improves the performance by up to compared to the existing inter-PE communication implementations. The implementation of PIDComm is available at https://github.com/AIS-SNU/PID-Comm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 795eafb4-6e03-4024-89e7-e7cefa5a7c88Cited by top-tier papers6
- Turbocharge ANNS on Real Processing-in-Memory by Enabling Fine-Grained Per-PIM-Core SchedulingPuqing Wu, Minhui Xie, Enrui Zhao, Dafang Zhang et al.USENIX ATC 2025 · 8 citations
- PIMnet: A Domain-Specific Network for Efficient Collective Communication in Scalable PIMHyojun Son, Gilbert Jonatan, Xiangyu Wu, Haeyoon Cho et al.HPCA 2025 · 7 citations
- DRIM-ANN: An Approximate Nearest Neighbor Search Engine based on Commercial DRAM-PIMsMingkai Chen, Tianhua Han, Cheng Liu, Shengwen Liang et al.SC 2025 · 5 citations
- Piccolo: Large-Scale Graph Processing with Fine-Grained in-Memory Scatter-GatherChangmin Shin, Jaeyong Song, Hongsun Jang, Dogeun Kim et al.HPCA 2025 · 5 citations
- DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory ArchitecturesPeiming Yang, Sankeerth Durvasula, Ivan Fernandez, Mohammad Sadrosadati et al.ISCA 2026 · 3 citations
Builds on22
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- RecNMP: Accelerating Personalized Recommendation with Near-Memory ProcessingLiu Ke, Udit Gupta, Benjamin Youngjae Cho, David Brooks et al.ISCA 2020 · 235 citations
- SANCUS: Staleness-Aware Communication-Avoiding Full-Graph Decentralized Training in Large-Scale Graph Neural NetworksJingshu Peng, Zhao Chen, Yingxia Shao, Yanyan Shen et al.VLDB 2022 · 76 citations
- TRiM: Enhancing Processor-Memory Interfaces with Scalable Tensor Reduction in MemoryJaehyun Park, Byeongho Kim, Sungmin Yun, Eojin Lee et al.MICRO 2021 · 70 citations
- Reducing communication in graph neural network trainingAlok Tripathy, Katherine A. Yelick, Aydin BuluçSC 2020 · 67 citations
Related papers
- Design and Analysis of a Processing-in-DIMM Join Algorithm: A Case Study with UPMEM DIMMsChaemin Lim, Suhyun Lee, Jinwoo Choi, Jounghoo Lee et al.SIGMOD 2023 · 50 citations
- CoCoTree: A Computation-Capable Architecture for Collective Communication in Scalable PIMShunchen Shi, Qijia Yang, Fan Yang, Yu Huang et al.HPCA 2026
- SPID-Join: A Skew-resistant Processing-in-DIMM Join Algorithm Exploiting the Bank- and Rank-level Parallelisms of DIMMsSuhyun Lee, Chaemin Lim, Jinwoo Choi, Heelim Choi et al.SIGMOD 2025 · 6 citations
- Accelerating Transactional Execution via Processing-In-MemoryAndré Lopes, Daniel Castro, Paolo RomanoEuroSys 2026
- ComPASS: A Compatible PIM Protocol Architecture and Scheduling Solution for Processor-PIM CollaborationSeunghyuk Yu, Hyeonu Kim, Kyoungho Jeun, Sunyoung Hwang et al.MICRO 2025 · 4 citations
