Lune

MICRO2025顶会

PolymorPIC: Embedding Polymorphic Processing-in-Cache in RISC-V based Processor for Full-stack Efficient AI Inference

Cheng Zou, Ziling Wei, Jun Yan Lee, Chen Nie, Kang You, Zhezhi He

2025年份
4被引次数
2顶会引用

摘要

The growing demand for neural network (NN) driven applications in AIoT devices necessitates efficient matrix multiplication (MM) acceleration.While domain-specific accelerators (DSAs) for NN are widely used, their large area overhead of dedicated buffers and low reusability limit cost-effectiveness.Processing-in-cache (PIC) architectures address this by repurposing existing SRAMs in processor caches for MM computation, eliminating dedicated DSA areas while retaining programmability.Despite its potential, PIC designs often overlook fundamental system-level issues, e.g., compactness, programmability, coherence, and scheduling optimization.In this work, we introduce PolymorPIC, a polymorphic architecture designed to accelerate MM directly within the cache using a bit-serial computing pattern.First, we propose a reconfigurable and processor-safe PIC architecture based on homogeneous memory arrays (HMAs), programmed through a meticulously designed interfaces in our software stack.Next, to enable mode switch of cache between cache mode and PIC mode, we develop a coherence strategy that ensures rapid, flexible, and processor-safe PIC.Moreover, we conduct scheduling optimization to maximize the NN acceleration performance of PolymorPIC.Ultimately, the PolymorPIC architecture is implemented on a RISC-V-based system-on-chip (SoC) and successfully end-to-end verified on a validation platform with an operating system running.Evaluation results show that by only introducing 11.5% area overhead for a single-core Out-of-Order processor (BOOM) with 1MB cache, PolymorPIC can improve the energy efficiency (TOPS/W) of multiple NNs by 1543.8× on average.Compared to system-level implementation of using NPU as co-processor (Gemmini), PolymorPIC outperforms it by 3.76× in area efficiency and 3.9× in energy efficiency.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖