Neuro-C: Neural Inference Shaped by Hardware Limits
Diletta Romano, Luca Mottola, Thiemo Voigt
Abstract
We present Neuro-centric Networks (Neuro-C), a neural network architecture we design to eliminate multiply-accumulate operations for efficient inference on ultra-low-power microcontrollers (MCUs). Although some MCUs include specialized hardware for neural acceleration, many ultra-low-power MCUs do not, requiring neural networks to align with limited compute and memory resources. Rather than compressing existing models or assuming dedicated hardware, Neuro-C integrates hardware constraints directly into the architecture, effectively shaping the network design around the limitations of the target platform. We shift the computational burden from connections to neurons and encode connectivity with a fixed ternary adjacency matrix, overcoming the bottleneck of matrix multiplications and large weight storage. This design enables a specialized inference kernel implementation that reduces memory usage and latency through pointer-based traversal and sparse dynamic memory allocation, complex control flows, and index decoding logic common in sparse or compressed models. Experimental results show that Neuro-C achieves accuracy comparable to or better than standard multilayer perceptrons across multiple datasets, while reducing inference latency and program memory usage by up to 90%. Compared to conventional ternary neural networks, Neuro-C provides improved convergence and accuracy under identical architectural settings, with negligible impact on inference latency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64acd16a-c8da-4478-a482-ce727e4c72b7Builds on1
Related papers
- Differentiable Neural Network Pruning to Enable Smart Applications on MicrocontrollersEdgar Liberis, Nicholas D. LaneUbiComp 2023 · 27 citations
- A2Q: Accumulator-Aware Quantization with Guaranteed Overflow AvoidanceIan Colbert, Alessandro Pappalardo, Jakoba Petri-KoenigICCV 2023 · 19 citations
- Intermittent Inference with Nonuniformly Compressed Multi-Exit Neural Network for Energy Harvesting Powered DevicesYawen Wu, Zhepeng Wang, Zhenge Jia, Yiyu Shi et al.DAC 2020 · 41 citations
- AdaptQNet: Optimizing Quantized DNN on Microcontrollers via Adaptive Heterogeneous Processing Unit UtilizationYansong Sun, Jialuo He, Dirk Kutscher, Huangxun ChenMobiCom 2025 · 1 citation
- NCPU: An Embedded Neural CPU Architecture on Resource-Constrained Low Power Devices for Real-time End-to-End PerformanceTianyu Jia, Yuhao Ju, Russ Joseph, Jie GuMICRO 2020 · 21 citations
