MASI: Memory-Adaptive Inference Framework for Spiking Neural Networks on Edge Devices
Di Yu, Helin Zheng, Changze Lv, Xin Du, Linshan Jiang, Xiang Liu, Gang Pan, Shuiguang Deng
Abstract
The rapid development of the Internet of Things (IoT) applications necessitates resource-efficient computing paradigms that can unify heterogeneous sensing modalities. Spiking Neural Networks (SNNs) meet this need with their event-driven and energy-efficient processing nature. However, deploying SNNs on mobile and embedded platforms is hindered by strict and fluctuating memory budgets. While prior work explores lightweight model design and system-level memory management, these methods either sacrifice accuracy or incur high runtime overhead due to timestep-dependent dynamics. To tackle these challenges, we propose a memory-adaptive framework MASI that enables efficient on-device SNN inference by combining (1) a fine-grained memory-adaptive layer slicing strategy, (2) a timestep-agnostic scheduler that maximizes memory utilization with minimal fragmentation, and (3) a timestep-aware early-exit mechanism that reduces redundant calculations. Evaluated on diverse workloads and edge devices, MASI can dynamically adapt to runtime memory availability, approximately reducing memory usage by 20.67% and inference latency by 58.53% on average with negligible accuracy loss compared to other feasible on-device implementations under memory constraints.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6bb3d1bf-a753-4eda-b0d1-2694ee5d125fRelated papers
- SpikeDyn: A Framework for Energy-Efficient Spiking Neural Networks with Continual and Unsupervised Learning Capabilities in Dynamic EnvironmentsRachmad Vidya Wicaksana Putra, Muhammad ShafiqueDAC 2021 · 4 citations
- FlexNN: Efficient and Adaptive DNN Inference on Memory-Constrained Edge DevicesXiangyu Li, Yuanchun Li, Yuanzhe Li, Ting Cao et al.MobiCom 2024 · 40 citations
- QP-SNN: Quantized and Pruned Spiking Neural NetworksWenjie Wei, Malu Zhang, Zijian Zhou, Ammar Belatreche et al.ICLR 2025
- Q-SNNs: Quantized Spiking Neural NetworksWenjie Wei, Yu Liang, Ammar Belatreche, Yichen Xiao et al.ACM MM 2024 · 23 citations
- Input-Aware Dynamic Timestep Spiking Neural Networks for Efficient In-Memory ComputingYuhang Li, Abhishek Moitra, Tamar Geller, Priyadarshini PandaDAC 2023 · 17 citations
