Neuro-C: Neural Inference Shaped by Hardware Limits
Diletta Romano, Luca Mottola, Thiemo Voigt
摘要
We present Neuro-centric Networks (Neuro-C), a neural network architecture we design to eliminate multiply-accumulate operations for efficient inference on ultra-low-power microcontrollers (MCUs). Although some MCUs include specialized hardware for neural acceleration, many ultra-low-power MCUs do not, requiring neural networks to align with limited compute and memory resources. Rather than compressing existing models or assuming dedicated hardware, Neuro-C integrates hardware constraints directly into the architecture, effectively shaping the network design around the limitations of the target platform. We shift the computational burden from connections to neurons and encode connectivity with a fixed ternary adjacency matrix, overcoming the bottleneck of matrix multiplications and large weight storage. This design enables a specialized inference kernel implementation that reduces memory usage and latency through pointer-based traversal and sparse dynamic memory allocation, complex control flows, and index decoding logic common in sparse or compressed models. Experimental results show that Neuro-C achieves accuracy comparable to or better than standard multilayer perceptrons across multiple datasets, while reducing inference latency and program memory usage by up to 90%. Compared to conventional ternary neural networks, Neuro-C provides improved convergence and accuracy under identical architectural settings, with negligible impact on inference latency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- Differentiable Neural Network Pruning to Enable Smart Applications on MicrocontrollersEdgar Liberis, Nicholas D. LaneUbiComp 2023 · 被引用 27 次
- A2Q: Accumulator-Aware Quantization with Guaranteed Overflow AvoidanceIan Colbert, Alessandro Pappalardo, Jakoba Petri-KoenigICCV 2023 · 被引用 19 次
- Intermittent Inference with Nonuniformly Compressed Multi-Exit Neural Network for Energy Harvesting Powered DevicesYawen Wu, Zhepeng Wang, Zhenge Jia, Yiyu Shi 等DAC 2020 · 被引用 41 次
- AdaptQNet: Optimizing Quantized DNN on Microcontrollers via Adaptive Heterogeneous Processing Unit UtilizationYansong Sun, Jialuo He, Dirk Kutscher, Huangxun ChenMobiCom 2025 · 被引用 1 次
- NCPU: An Embedded Neural CPU Architecture on Resource-Constrained Low Power Devices for Real-time End-to-End PerformanceTianyu Jia, Yuhao Ju, Russ Joseph, Jie GuMICRO 2020 · 被引用 21 次
