NCPU: An Embedded Neural CPU Architecture on Resource-Constrained Low Power Devices for Real-time End-to-End Performance
Tianyu Jia, Yuhao Ju, Russ Joseph, Jie Gu
Abstract
Machine learning inference has become an essential task for embedded edge devices requiring the deployment of costly deep neural network accelerators onto extremely resource-constrained hardware. Although many optimization strategies have been proposed to improve the efficiency of standalone accelerators, the optimization for end-to-end performance of a computing device with heterogeneous cores is still challenging and often overlooked, especially for low power devices. In this paper, we propose a unified reconfigurable architecture, referred as Neural CPU (NCPU), for low-cost embedded systems. The proposed architecture is built on a binary neural network accelerator with the capability to emulate an in-order RISC-V CPU pipeline. The NCPU supports flexible programmability of RISC-V and maintains data locally to avoid costly core-to-core data transfer. A two-core NCPU SoC is designed and fabricated in a 65nm CMOS process. Compared with the conventional heterogeneous architecture, a single NCPU achieves 35% area reduction and 12% energy saving at 0.4V, which is suitable for low power and low-cost embedded edge devices. The NCPU design also features the capability of smooth switching between general-purpose CPU operation and a binary neural network inference to realize full utilization of the cores. The implemented two-core NCPU SoC achieves an end-to-end performance speed-up of 43% or an equivalent 74% energy saving based on use cases of real-time image classification and motion detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9819405e-da32-4abe-83ee-3570e4f8febeRelated papers
- BiSon-e: a lightweight and high-performance accelerator for narrow integer linear algebra computing on the edgeEnrico Reggiani, Cristóbal Ramírez Lazo, Roger Figueras Bagué, Adrián Cristal et al.ASPLOS 2022 · 11 citations
- Flex: Fast, Accurate DNN Inference on Low-Cost Edges Using Heterogeneous Accelerator ExecutionTanmoy Sen, Haiying Shen, Anand Padmanabha IyerEuroSys 2025 · 2 citations
- EdgeNN: Efficient Neural Network Inference for CPU-GPU Integrated Edge DevicesChenyang Zhang, Feng Zhang, Kuangyu Chen, Mingjun Chen et al.ICDE 2023 · 16 citations
- Parallel DNN Inference Framework Leveraging a Compact RISC-V ISA-based Multi-core SystemYipeng Zhang, Bo Du, Lefei Zhang, Jia WuKDD 2020 · 14 citations
- Mix-GEMM: An efficient HW-SW Architecture for Mixed-Precision Quantized Deep Neural Networks Inference on Edge DevicesEnrico Reggiani, Alessandro Pappalardo, Max Doblas, Miquel Moretó et al.HPCA 2023 · 31 citations
