PolymorPIC: Embedding Polymorphic Processing-in-Cache in RISC-V based Processor for Full-stack Efficient AI Inference
Cheng Zou, Ziling Wei, Jun Yan Lee, Chen Nie, Kang You, Zhezhi He
摘要
The growing demand for neural network (NN) driven applications in AIoT devices necessitates efficient matrix multiplication (MM) acceleration.While domain-specific accelerators (DSAs) for NN are widely used, their large area overhead of dedicated buffers and low reusability limit cost-effectiveness.Processing-in-cache (PIC) architectures address this by repurposing existing SRAMs in processor caches for MM computation, eliminating dedicated DSA areas while retaining programmability.Despite its potential, PIC designs often overlook fundamental system-level issues, e.g., compactness, programmability, coherence, and scheduling optimization.In this work, we introduce PolymorPIC, a polymorphic architecture designed to accelerate MM directly within the cache using a bit-serial computing pattern.First, we propose a reconfigurable and processor-safe PIC architecture based on homogeneous memory arrays (HMAs), programmed through a meticulously designed interfaces in our software stack.Next, to enable mode switch of cache between cache mode and PIC mode, we develop a coherence strategy that ensures rapid, flexible, and processor-safe PIC.Moreover, we conduct scheduling optimization to maximize the NN acceleration performance of PolymorPIC.Ultimately, the PolymorPIC architecture is implemented on a RISC-V-based system-on-chip (SoC) and successfully end-to-end verified on a validation platform with an operating system running.Evaluation results show that by only introducing 11.5% area overhead for a single-core Out-of-Order processor (BOOM) with 1MB cache, PolymorPIC can improve the energy efficiency (TOPS/W) of multiple NNs by 1543.8× on average.Compared to system-level implementation of using NPU as co-processor (Gemmini), PolymorPIC outperforms it by 3.76× in area efficiency and 3.9× in energy efficiency.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- NasZip: Software and Hardware Co-Design to Accelerate Approximate Nearest Neighbor Search with DIMM-Based Near-Data ProcessingCheng Zou, Shuo Yang, Chen Nie, Yu Zou 等ISCA 2026 · 被引用 1 次
- ELSA: An Elastic Snn Inference Architecture for Efficient Neuromorphic ComputingKang You, Chen Nie, Lee Jun Yan, Ziling Wei 等ISCA 2026
相关 Paper
- Parallel DNN Inference Framework Leveraging a Compact RISC-V ISA-based Multi-core SystemYipeng Zhang, Bo Du, Lefei Zhang, Jia WuKDD 2020 · 被引用 14 次
- MAICC : A Lightweight Many-core Architecture with In-Cache Computing for Multi-DNN Parallel InferenceRenhao Fan, Yikai Cui, Qilin Chen, Mingyu Wang 等MICRO 2023 · 被引用 12 次
- NCPU: An Embedded Neural CPU Architecture on Resource-Constrained Low Power Devices for Real-time End-to-End PerformanceTianyu Jia, Yuhao Ju, Russ Joseph, Jie GuMICRO 2020 · 被引用 21 次
- DataMaestro: A Versatile and Efficient Data Streaming Engine Bringing Decoupled Memory Access To Dataflow AcceleratorsXiaoling Yi, Yunhao Deng, Ryan Antonio, Fanchen Kong 等DAC 2025 · 被引用 4 次
- RAELLA: Reforming the Arithmetic for Efficient, Low-Resolution, and Low-Loss Analog PIM: No Retraining Required!Tanner Andrulis, Joel S. Emer, Vivienne SzeISCA 2023 · 被引用 45 次
