GradPIM: A Practical Processing-in-DRAM Architecture for Gradient Descent
Heesu Kim, Hanmin Park, Taehyun Kim, Kwanheum Cho, Eojin Lee, Soojung Ryu, Hyuk-Jae Lee, Kiyoung Choi, Jinho Lee
Abstract
In this paper, we present GradPIM, a processingin-memory architecture which accelerates parameter updates of deep neural networks training. As one of processing-in-memory techniques that could be realized in the near future, we propose an incremental, simple architectural design that does not invade the existing memory protocol. Extending DDR4 SDRAM to utilize bank-group parallelism makes our operation designs in processing-in-memory (PIM) module efficient in terms of hardware cost and performance. Our experimental results show that the proposed architecture can improve the performance of DNN training and greatly reduce memory bandwidth requirement while posing only a minimal amount of overhead to the protocol and DRAM area.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6714e42d-362d-4472-9ac6-13b73b89aebcCited by top-tier papers13
- NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM InferencingGuseul Heo, Sangyeop Lee, Jaehong Cho, Hyunmin Choi et al.ASPLOS 2024 · 121 citations
- Pathfinding Future PIM Architectures by Demystifying a Commercial PIM TechnologyBongjoon Hyun, Taehun Kim, Dongjae Lee, Minsoo RhuHPCA 2024 · 62 citations
- Flash-Cosmos: In-Flash Bulk Bitwise Operations Using Inherent Computation Capability of NAND Flash MemoryJisung Park, Roknoddin Azizi, Geraldo F. Oliveira, Mohammad Sadrosadati et al.MICRO 2022 · 53 citations
- Sparse Attention Acceleration with Synergistic In-Memory Pruning and On-Chip RecomputationAmir Yazdanbakhsh, Ashkan Moradifirouzabadi, Zheng Li, Mingu KangMICRO 2022 · 47 citations
- MeNDA: a near-memory multi-way merge solution for sparse transposition and dataflowsSiying Feng, Xin He, Kuan-Yu Chen, Liu Ke et al.ISCA 2022 · 28 citations
Related papers
- SplitSync: Bank Group-Level Split-Synchronization for High-Performance DRAM PIMByungkuk Yoon, Sanghyeok Han, Gyeonghwan Park, Jae-Joon KimDAC 2025 · 1 citation
- GoPIM: GCN-Oriented Pipeline Optimization for PIM AcceleratorsSiling Yang, Shuibing He, Wenjiong Wang, Yanlong Yin et al.HPCA 2025 · 3 citations
- INCA: Input-stationary Dataflow at Outside-the-box Thinking about Deep Learning AcceleratorsBokyung Kim, Shiyu Li, Hai LiHPCA 2023 · 28 citations
- UpDLRM: Accelerating Personalized Recommendation using Real-World PIM ArchitectureSitian Chen, Haobin Tan, Amelie Chi Zhou, Yusen Li et al.DAC 2024 · 9 citations
- Accelerating Aggregation Using a Real Processing-in-Memory SystemMuhammad Attahir Jibril, Hani Al-Sayeh, Kai-Uwe SattlerICDE 2024 · 7 citations
