HH-PIM: Dynamic Optimization of Power and Performance with Heterogeneous-Hybrid PIM for Edge AI Devices
Sangmin Jeon, Kangju Lee, Kyeongwon Lee, Woojoo Lee
Abstract
Processing-in-Memory (PIM) architectures offer promising solutions for efficiently handling AI applications in energy-constrained edge environments. While traditional PIM designs enhance performance and energy efficiency by reducing data movement between memory and processing units, they are limited in edge devices due to continuous power demands and the storage requirements of large neural network weights in SRAM and DRAM. Hybrid PIM architectures, incorporating nonvolatile memories like MRAM and ReRAM, mitigate these limitations but struggle with a mismatch between fixed computing resources and dynamically changing inference workloads. To address these challenges, this study introduces a Heterogeneous-Hybrid PIM (HH-PIM) architecture, comprising high-performance MRAM-SRAM PIM modules and low-power MRAM-SRAM PIM modules. We further propose a data placement optimization algorithm that dynamically allocates data based on computational demand, maximizing energy efficiency. FPGA prototyping and power simulations with processors featuring HH-PIM and other PIM types demonstrate that the proposed HH-PIM achieves up to 60.43% average energy savings over conventional PIMs while meeting application latency requirements. These results confirm HH-PIM’s suitability for adaptive, energy-efficient AI processing in edge devices.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8b4f2be0-d26a-4fa7-8290-d3396b2d2e4fBuilds on5
- PIMGCN: A ReRAM-Based PIM Design for Graph Convolutional Network AccelerationTao Yang, Dongyue Li, Yibo Han, Yilong Zhao et al.DAC 2021 · 39 citations
- Accelerating Sparse Attention with a Reconfigurable Non-volatile Processing-In-Memory ArchitectureQilin Zheng, Shiyu Li, Yitu Wang, Ziru Li et al.DAC 2023 · 14 citations
- A Model-Specific End-to-End Design Methodology for Resource-Constrained TinyML HardwareYanchi Dong, Tianyu Jia, Kaixuan Du, Yiqi Jing et al.DAC 2023 · 10 citations
- PIM-HLS: An Automatic Hardware Generation Tool for Heterogeneous Processing-In-Memory-based Neural Network AcceleratorsYu Zhu, Zhenhua Zhu, Guohao Dai, Fengbin Tu et al.DAC 2023 · 9 citations
- Efficient Memory Integration: MRAM-SRAM Hybrid Accelerator for Sparse On-Device LearningFan Zhang, Amitesh Sridharan, Wilman Tsai, Yiran Chen et al.DAC 2024 · 7 citations
Related papers
- PIM-MMU: A Memory Management Unit for Accelerating Data Transfers in Commercial PIM SystemsDongjae Lee, Bongjoon Hyun, Taehun Kim, Minsoo RhuMICRO 2024 · 23 citations
- FACIL: Flexible DRAM Address Mapping for SoC-PIM Cooperative On-device LLM InferenceSeong Hoon Seo, Junghoon Kim, Donghyun Lee, Seonah Yoo et al.HPCA 2025 · 7 citations
- Processing-in-SRAM acceleration for ultra-low power visual 3D perceptionYuquan He, Songyun Qu, Gangliang Lin, Cheng Liu et al.DAC 2022 · 4 citations
- CREAM: computing in ReRAM-assisted energy and area-efficient SRAM for neural network accelerationLiukai Xu, Songyuan Liu, Zhi Li, Dengfeng Wang et al.DAC 2022 · 6 citations
- On Endurance of Processing in (Nonvolatile) MemorySalonik Resch, M. Hüsrev Cilasun, Zamshed I. Chowdhury, Masoud Zabihi et al.ISCA 2023 · 15 citations
