PipeIMC: A Pipelined In-SRAM Computing Architecture
Yikai Cui, Renhao Fan, Weike Li, Mingzhao Li, Mingyu Wang, Zhaolin Li
摘要
Generative large language models pose a significant challenge to the memory bandwidth and computing capabilities of traditional computing architectures. By performing computations inside the SRAM arrays, in-SRAM computing architectures can achieve substantial performance while reducing memory hierarchy bandwidth usage and energy consumption, making them a strong alternative to CPUs and GPUs for executing LLMs. However, prior in-SRAM computing architectures have found it challenging to achieve higher performance, because they are restricted by their in-order execution mechanism, in which each operation must wait for the completion of its preceding operations before execution. To address this problem, this paper proposes PipeIMC, a pipelined in-SRAM computing architecture with two stages: memory and calculation. Accordingly, each in-SRAM computing operation is segmented into the memory phase and the calculation phase. When the calculation phase of the current operation is being executed in the calculation stage, the memory stage can execute the memory phase of the next operation to fetch the required data. To alleviate the influence of data hazards and control hazards, we propose an out-of-order execution mechanism integrated with explicit register renaming. To further enhance performance, we propose a fine-grained issue mechanism that enables the next operation to be issued earlier during the idle cycles of the memory stage of the current operation. Evaluation results show that PipeIMC achieves 2.15x to 3.96x and 1.13x to 4.77x utilization, compared to EVE and Duality Cache, two state-of-the-art in-SRAM computing architectures. This improvement in utilization yields a performance of 2.17x and 1.68x per area, and an energy efficiency of 1.92 x and 1.60 x, on average, over EVE and Duality Cache on the Rodinia GPU benchmarks.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Ouroboros: Wafer-Scale SRAM CIM with Token-Grained Pipelining for Large Language Model InferenceYiqi Liu, Yudong Pan, Mengdi Wang, Shixin Zhao 等ASPLOS 2026 · 被引用 1 次
- Bridging Efficiency and Scalability in Llm System Via 3D Hybrid Pim With 2D in-Transit ComputationHongyi Li, Songchen Ma, Huanyu Qu, Weihao Zhang 等ISCA 2026
- PIMPAL: Accelerating LLM Inference on Edge Devices via In-DRAM Arithmetic LookupYoonho Jang, Hyeongjun Cho, Yesin Ryu, Jungrae Kim 等DAC 2025 · 被引用 6 次
- IANUS: Integrated Accelerator based on NPU-PIM Unified Memory SystemMinseok Seo, Xuan Truong Nguyen, Seok Joong Hwang, Yongkee Kwon 等ASPLOS 2024 · 被引用 57 次
- A Full-system, Programmable, and Extensible In-Memory Computing Simulation Framework for Deep LearningKaining Zhou, Jian Huang, Nam Sung Kim, Naresh ShanbhagDAC 2025
