SC2021Top-tier venue
Accelerating bandwidth-bound deep learning inference with main-memory accelerators
Benjamin Y. Cho, Jeageun Jung, Mattan Erez
Abstract
DL inference queries play an important role in diverse internet services and a large fraction of datacenter cycles are spent on processing DL inference queries. Specifically, the matrixmatrix multiplication (GEMM) operations of fully-connected MLP layers dominate many inference tasks. We find that the GEMM operations for datacenter DL inference tasks are memory bandwidth bound, contrary to common assumptions: (1) strict query latency constraints force small-batch operation, which limits reuse and increases bandwidth demands; and (2) large and colocated models require reading the large weight matrices from main memory, again requiring high bandwidth without offering reuse opportunities. We demonstrate the large potential of accelerating these small-batch GEMMs with processing in the main CPU memory. We develop a novel GEMM execution flow and corresponding memory-side address-generation logic that exploits GEMM locality and enables long-running PIM kernels despite the complex address-mapping functions employed by the CPU that would otherwise destroy locality. Our evaluation of StepStone variants at the channel, device, and within-device PIM levels, along with optimizations that balance parallelism benefits with data-distribution overheads demonstrate 12× better minimum latency than a CPU and 2.8× greater throughput for strict query latency constraints. End-to-end performance analysis of recent recommendation and language models shows that StepStone PIM outperforms a fast CPU (by up to 16×) and prior main-memory acceleration approaches (by up to 2.4× compared to the best prior approach).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM InferencingGuseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi et al.ASPLOS 2024 · 121 citations
- Pathfinding Future PIM Architectures by Demystifying a Commercial PIM TechnologyBongjoon Hyun, Taehun Kim, Dongjae Lee, Minsoo RhuHPCA 2024 · 62 citations
- MeNDA: a near-memory multi-way merge solution for sparse transposition and dataflowsSiying Feng, Xin He, Kuan-Yu Chen, Liu Ke et al.ISCA 2022 · 28 citations
- PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM SystemsDongjae Lee, Bongjoon Hyun, Taehun Kim, Minsoo RhuMICRO 2024 · 23 citations
- Turbocharge ANNS on Real Processing-in-Memory by Enabling Fine-Grained Per-PIM-Core SchedulingPuqing Wu, Minhui Xie, Enrui Zhao, Dafang Zhang et al.USENIX ATC 2025 · 8 citations
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- DRAMA: Exploiting DRAM Addressing for Cross-CPU AttacksPeter Pessl, Daniel Gruss, Clémentine Maurice, Michael Schwarz et al.USENIX Security 2016 · 500 citations
- RecNMP: Accelerating Personalized Recommendation with Near-Memory ProcessingLiu Ke, Udit Gupta, Benjamin Youngjae Cho, David Brooks et al.ISCA 2020 · 235 citations
- Newton: A DRAM-maker's Accelerator-in-Memory (AiM) Architecture for Machine LearningMingxuan He, Choungki Song, Ilkon Kim, Chunseok Jeong et al.MICRO 2020 · 208 citations
- DeepRecSys: A System for Optimizing End-To-End At-Scale Neural Recommendation InferenceUdit Gupta, Samuel Hsia, Vikram Saraph, Xiaodong Wang et al.ISCA 2020 · 149 citations
Related papers
- UpDLRM: Accelerating Personalized Recommendation using Real-World PIM ArchitectureSitian Chen, Haobin Tan, Amelie Chi Zhou, Yusen Li et al.DAC 2024 · 9 citations
- Optimizing CPU Performance for Recommendation Systems At-ScaleRishabh Jain, Scott Cheng, Vishwas Kalagi, Vrushabh Sanghavi et al.ISCA 2023 · 25 citations
- DECA: A Near-Core LLM Decompression Accelerator Grounded on a 3D Roofline ModelGerasimos Gerogiannis, Stijn Eyerman, Evangelos Georganas, Wim Heirman et al.MICRO 2025 · 5 citations
- BlockPIM: Optimizing Memory Management for PIM-enabled Long-Context LLM InferenceZhichun Li, Jun Zhou, Xueqi Li, Ninghui SunDAC 2025 · 3 citations
- RecFlow: Unlocking GPU Efficiency for DLRM Inference via Fine-Grained Parallelism and Incremental BatchingSiheng Pan, Shaolong Li, Minwei Zhang, Shuxi Guo et al.INFOCOM 2026
