ABC-DIMM: Alleviating the Bottleneck of Communication in DIMM-based Near-Memory Processing with Inter-DIMM Broadcast
Weiyi Sun, Zhaoshi Li, Shouyi Yin, Shaojun Wei, Leibo Liu
摘要
Near-Memory Processing (NMP) systems that integrate accelerators within DIMM (Dual-Inline Memory Module) buffer chips potentially provide high performance with relatively low design and manufacturing costs. However, an inevitable communication bottleneck arises when considering the main memory bus among peer DIMMs and the host CPU. This communication bottleneck roots in the bus-based nature and the limited point-to-point communication pattern of the main memory system. The aggregated memory bandwidth of DIMM- based NMP scales with the number of DIMMs. When the number of DIMMs in a channel scales up, the per-DIMM point-to-point communication bandwidth scales down, whereas the computation resources and local memory bandwidth per DIMM stay the same. For many important sparse data-intensive workloads like graph applications and sparse tensor algebra, we identify that communication among DIMMs and the host CPU easily dominates their processing procedure in previous DIMM-based NMP systems, which severely bottlenecks their performance.To tackle this challenge, we propose that inter-DIMM broadcast should be implemented and utilized in the main memory system of DIMM-based NMP. On the hardware side, the main memory bus naturally scales out with broadcast, where per- DIMM effective bandwidth of broadcast remains the same as the number of DIMMs grows. On the software side, many sparse applications can be implemented in a form such that broadcasts dominate their communication. Based on these ideas, we design ABC-DIMM, which Alleviates the Bottleneck of Communication in DIMM-based NMP, consisting of integral broadcast mechanisms and Broadcast-Process programming framework, with minimized modifications to commodity software-hardware stack. Our evaluation shows that ABC-DIMM offers an 8.33 × geo-mean speedup over a 16-core CPU baseline, and outperforms two NMP baselines by 2.59 × and 2.93 × on average.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper13
- DIMMining: pruning-efficient and parallel graph mining on near-memory-computingGuohao Dai, Zhenhua Zhu, Tianyu Fu, Chiyue Wei 等ISCA 2022 · 被引用 56 次
- ABNDP: Co-optimizing Data Access and Load Balance in Near-Data ProcessingBoyu Tian, Qihang Chen, Mingyu GaoASPLOS 2023 · 被引用 31 次
- MeNDA: a near-memory multi-way merge solution for sparse transposition and dataflowsSiying Feng, Xin He, Kuan-Yu Chen, Liu Ke 等ISCA 2022 · 被引用 28 次
- NDPBridge: Enabling Cross-Bank Coordination in Near-DRAM-Bank Processing ArchitecturesBoyu Tian, Yiwei Li, Li Jiang, Shuangyu Cai 等ISCA 2024 · 被引用 27 次
- Low-Overhead General-Purpose Near-Data Processing in CXL Memory ExpandersHyungkyu Ham, Jeongmin Hong, Geonwoo Park, Yunseon Shin 等MICRO 2024 · 被引用 26 次
相关 Paper
- DIMM-Link: Enabling Efficient Inter-DIMM Communication for Near-Memory ProcessingZhe Zhou, Cong Li, Fan Yang, Guangyu SunHPCA 2023 · 被引用 42 次
- MetaNMP: Leveraging Cartesian-Like Product to Accelerate HGNNs with Near-Memory ProcessingDan Chen, Haiheng He, Hai Jin, Long Zheng 等ISCA 2023 · 被引用 36 次
- AsyncDIMM: Achieving Asynchronous Execution in DIMM-Based Near-Memory ProcessingLiyan Chen, Dongxu Lyu, Jianfei Jiang, Qin Wang 等HPCA 2025 · 被引用 7 次
- NMExplorer: An Efficient Exploration Framework for DIMM-based Near-Memory Tensor ReductionCong Li, Zhe Zhou, Xingchen Li, Guangyu Sun 等DAC 2023 · 被引用 5 次
- Design and Analysis of a Processing-in-DIMM Join Algorithm: A Case Study with UPMEM DIMMsChaemin Lim, Suhyun Lee, Jinwoo Choi, Jounghoo Lee 等SIGMOD 2023 · 被引用 50 次
