AsyncDIMM: Achieving Asynchronous Execution in DIMM-Based Near-Memory Processing
Liyan Chen, Dongxu Lyu, Jianfei Jiang, Qin Wang, Zhigang Mao, Naifeng Jing
Abstract
DIMM-based near-memory processing (NMP) architectures address the “memory wall” problem by incorporating near-memory accelerators (NMAs) into main memory devices for high memory bandwidth and low energy consumption. However, critical challenges prevent efficient asynchronous execution between host and NMAs in DIMM-NMP architectures. Memory controllers (MCs) distributed at the host side and the NMA side issue memory accesses independently without synchronization on memory states, which may lead to memory bus contention and DRAM errors. Therefore, most existing DIMM-NMP designs adopt synchronous execution to prevent concurrent memory accesses. However, this intervention wastes either the host or the NMA computation capability. In this work, we propose AsyncDIMM, a novel DIMM-NMP design with efficient asynchronous execution based on existing memory buses. It enables single access mode (host or NMA), concurrent access mode, and a seamless switch between them. First, we propose the offload-schedule-return mechanism with explicit and implicit synchronization to ensure memory access correctness for all memory modes. Second, to further improve bandwidth utilization and decrease access latency, we introduce optimized timing constraints for offloading, a locality-aware switch-recovery method for scheduling, and adaptive batch with timing-division multiplexing notification for returning. Finally, we present a detailed design with limited hardware modifications to conventional host and NMA MCs, which is extensively validated on the FPGA. Comprehensive experiments demonstrate that AsyncDIMM outperforms four NMP baselines by , enabling efficient asynchronous execution with up to bandwidth utilization uplift and 47% access latency reduction.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers4
- No Cap, This Memory Slaps: Breaking Through the Memory Wall of Transactional Database Systems with Processing-in-MemoryHyoungjoo Kim, Yiwei Zhao, Andrew Pavlo, Phillip B. GibbonsVLDB 2025 · 7 citations
- CHIME: A Case for Efficient Long-Context Attention-FC Disaggregated Inference with DIMM-PIMQingyuan Liu, Liyan Chen, Haocheng Wang, Yanning Yang et al.ISCA 2026 · 6 citations
- COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile DevicesYilong Zhao, Fangxin Liu, Onur Mutlu, Mingyu Gao et al.ISCA 2026 · 1 citation
- LoCaLUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIMJunguk Hong, Changmin Shin, Sukjin Kim, Si Ung Noh et al.HPCA 2026
Related papers
- DIMM-Link: Enabling Efficient Inter-DIMM Communication for Near-Memory ProcessingZhe Zhou, Cong Li, Fan Yang, Guangyu SunHPCA 2023 · 42 citations
- ABC-DIMM: Alleviating the Bottleneck of Communication in DIMM-based Near-Memory Processing with Inter-DIMM BroadcastWeiyi Sun, Zhaoshi Li, Shouyi Yin, Shaojun Wei et al.ISCA 2021 · 42 citations
- UniNDP: A Unified Compilation and Simulation Tool for Near DRAM Processing ArchitecturesTongxin Xie, Zhenhua Zhu, Bing Li, Yukai He et al.HPCA 2025 · 9 citations
- HAIL-DIMM: Host Access Interleaved with Near-Data Processing on DIMM-based Memory SystemMinkyu Lee, Sang-Seol Lee, Kyungho Kim, Eunchong Lee et al.DAC 2024 · 2 citations
- ABNDP: Co-optimizing Data Access and Load Balance in Near-Data ProcessingBoyu Tian, Qihang Chen, Mingyu GaoASPLOS 2023 · 31 citations
