PolymorPIC: Embedding Polymorphic Processing-in-Cache in RISC-V based Processor for Full-stack Efficient AI Inference
Cheng Zou, Ziling Wei, Jun Yan Lee, Chen Nie, Kang You, Zhezhi He
Abstract
The growing demand for neural network (NN) driven applications in AIoT devices necessitates efficient matrix multiplication (MM) acceleration.While domain-specific accelerators (DSAs) for NN are widely used, their large area overhead of dedicated buffers and low reusability limit cost-effectiveness.Processing-in-cache (PIC) architectures address this by repurposing existing SRAMs in processor caches for MM computation, eliminating dedicated DSA areas while retaining programmability.Despite its potential, PIC designs often overlook fundamental system-level issues, e.g., compactness, programmability, coherence, and scheduling optimization.In this work, we introduce PolymorPIC, a polymorphic architecture designed to accelerate MM directly within the cache using a bit-serial computing pattern.First, we propose a reconfigurable and processor-safe PIC architecture based on homogeneous memory arrays (HMAs), programmed through a meticulously designed interfaces in our software stack.Next, to enable mode switch of cache between cache mode and PIC mode, we develop a coherence strategy that ensures rapid, flexible, and processor-safe PIC.Moreover, we conduct scheduling optimization to maximize the NN acceleration performance of PolymorPIC.Ultimately, the PolymorPIC architecture is implemented on a RISC-V-based system-on-chip (SoC) and successfully end-to-end verified on a validation platform with an operating system running.Evaluation results show that by only introducing 11.5% area overhead for a single-core Out-of-Order processor (BOOM) with 1MB cache, PolymorPIC can improve the energy efficiency (TOPS/W) of multiple NNs by 1543.8× on average.Compared to system-level implementation of using NPU as co-processor (Gemmini), PolymorPIC outperforms it by 3.76× in area efficiency and 3.9× in energy efficiency.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3f685936-e050-4b9e-b660-34bdd24d043bCited by top-tier papers2
- NasZip: Software and Hardware Co-Design to Accelerate Approximate Nearest Neighbor Search with DIMM-Based Near-Data ProcessingCheng Zou, Shuo Yang, Chen Nie, Yu Zou et al.ISCA 2026 · 1 citation
- ELSA: An Elastic Snn Inference Architecture for Efficient Neuromorphic ComputingKang You, Chen Nie, Lee Jun Yan, Ziling Wei et al.ISCA 2026
Related papers
- Parallel DNN Inference Framework Leveraging a Compact RISC-V ISA-based Multi-core SystemYipeng Zhang, Bo Du, Lefei Zhang, Jia WuKDD 2020 · 14 citations
- MAICC : A Lightweight Many-core Architecture with In-Cache Computing for Multi-DNN Parallel InferenceRenhao Fan, Yikai Cui, Qilin Chen, Mingyu Wang et al.MICRO 2023 · 12 citations
- NCPU: An Embedded Neural CPU Architecture on Resource-Constrained Low Power Devices for Real-time End-to-End PerformanceTianyu Jia, Yuhao Ju, Russ Joseph, Jie GuMICRO 2020 · 21 citations
- DataMaestro: A Versatile and Efficient Data Streaming Engine Bringing Decoupled Memory Access To Dataflow AcceleratorsXiaoling Yi, Yunhao Deng, Ryan Antonio, Fanchen Kong et al.DAC 2025 · 4 citations
- RAELLA: Reforming the Arithmetic for Efficient, Low-Resolution, and Low-Loss Analog PIM: No Retraining Required!Tanner Andrulis, Joel S. Emer, Vivienne SzeISCA 2023 · 45 citations
