HEIRS: Hybrid Three-Dimension RRAM- and SRAM-CIM Architecture for Multi-task Transformer Acceleration
Liukai Xu, Shuai Yuan, Dengfeng Wang, Yiming Chen, Xueqing Li, Yanan Sun
Abstract
Large-scale transformer with millions of weights achieves great success in multiple natural language processing (NLP) tasks. To release the memory bottleneck of multi-task model deployment, transfer learning tunes part of weights with shared parameters among tasks. Moreover, computing-in-memory (CIM) emerges as an efficient solution for neural network acceleration. With higher storage density, RRAM-CIM can store the large-scale model without costly weight loading, compared with another mainstream SRAM-CIM. However, the RRAM rewrite for tuned and dynamic weight matrix-vector-multiplication (MVM) in transformers requires high-cost RRAM writing in RRAM-CIM. Current hybrid CIM can compensate for the weakness of RRAM-CIM by adding SRAM-CIM with independent MVM. However, the tuned weights in transfer learning cannot be implemented due to the demand for the cooperative addition of MVM results from both shared and tuned weights. In this paper, a hybrid three-dimension RRAM-CIM and SRAM-CIM architecture (HEIRS) is proposed for multi-task transformer acceleration, with monolithically 3D integration of high-density RRAM-CIM and high-performance SRAM-CIM. The 3D RRAM-CIM with ultra-high density stores the whole model with mitigated off-chip weight loading. The SRAM-CIM is employed for efficiently performing dynamic weight MVM without RRAM rewrite. Moreover, a novel hybrid-CIM paradigm is proposed with an input selective adder tree, to support cooperative addition in transfer learning. Experiments show that, compared with RRAM-CIM and SRAM-CIM, the proposed HEIRS improves the energy efficiency by up to 7.83x and 2.29x on BERT, respectively. Meanwhile, the latency is also reduced by up to 85.5% and the storage density is enhanced by 7.2x, compared to RRAM-CIM.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 5edba1de-c3ea-40db-acdf-8c879b4cc154Cited by top-tier papers2
- UniCAIM: A Unified CAM/CIM Architecture with Static-Dynamic KV Cache Pruning for Efficient Long-Context LLM InferenceWeikai Xu, Wenxuan Zeng, Qianqian Huang, Meng Li et al.DAC 2025 · 3 citations
- DARTH-PUM: A Hybrid Processing-Using-Memory ArchitectureRyan Wong, Ben Feinberg, Saugata GhoseASPLOS 2026 · 1 citation
Related papers
- INCA: Input-stationary Dataflow at Outside-the-box Thinking about Deep Learning AcceleratorsBokyung Kim, Shiyu Li, Hai LiHPCA 2023 · 28 citations
- CREAM: computing in ReRAM-assisted energy and area-efficient SRAM for neural network accelerationLiukai Xu, Songyuan Liu, Zhi Li, Dengfeng Wang et al.DAC 2022 · 6 citations
- Efficient Edge Vision Transformer Accelerator with Decoupled Chunk Attention and Hybrid Computing-In-MemoryYi Li, Zijian Ye, Xiangqu Fu, Songqi Wang et al.DAC 2025 · 2 citations
- Hybrid SLC-MLC RRAM Mixed-Signal Processing-in-Memory Architecture for Transformer Acceleration via Gradient RedistributionChang Eun Song, Priyansh Bhatnagar, Zihan Xia, Nam Sung Kim et al.ISCA 2025 · 4 citations
- Ouroboros: Wafer-Scale SRAM CIM with Token-Grained Pipelining for Large Language Model InferenceYiqi Liu, Yudong Pan, Mengdi Wang, Shixin Zhao et al.ASPLOS 2026 · 1 citation
