Boosting Mobile CNN Inference through Semantic Memory
Yun Li, Chen Zhang, Shihao Han, Li Lyna Zhang, Baoqun Yin, Yunxin Liu, Mengwei Xu
Abstract
Human brains are known to be capable of speeding up visual recognition of repeatedly presented objects through faster memory encoding and accessing procedures on activated neurons. For the first time, we borrow and distill such a capability into a semantic memory design, namely SMTM, to improve on-device CNN inference. SMTM employs a hierarchical memory architecture to leverage the long-tail distribution of objects of interest, and further incorporates several novel techniques to put it into effects: (1) it encodes high-dimensional feature maps into low-dimensional, semantic vectors for low-cost yet accurate cache and lookup; (2) it uses a novel metric in determining the exit timing considering different layers' inherent characteristics; (3) it adaptively adjusts the cache size and semantic vectors to fit the scene dynamics. SMTM is prototyped on commodity CNN engine and runs on both mobile CPU and GPU. Extensive experiments on large-scale datasets and models show that SMTM can significantly speed up the model inference over standard approach (up to 2×) and prior cache designs (up to 1.5x), with acceptable accuracy loss.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97fefdb4-9632-4841-9b1e-221d5764f795Cited by top-tier papers4
- LUT-NN: Empower Efficient Neural Network Inference with Centroid Learning and Table LookupXiaohu Tang, Yang Wang, Ting Cao, Li Lyna Zhang et al.MobiCom 2023 · 29 citations
- LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning AcceleratorGuoyu Li, Shengyu Ye, Chunyun Chen, Yang Wang et al.HPCA 2025 · 7 citations
- Accelerating End-Cloud Collaborative Inference via Near Bubble-Free Pipeline OptimizationLuyao Gao, Jianchun Liu, Hongli Xu, Sun Xu et al.INFOCOM 2025 · 4 citations
- Many Hands Make Light Work: Accelerating Edge Inference via Multi-Client Collaborative CachingWenyi Liang, Jianchun Liu, Hongli Xu, Chunming Qiao et al.ICDE 2025
Builds on6
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo et al.ICCV 2019 · 633 citations
- SPINN: synergistic progressive inference of neural networks over device and cloudStefanos Laskaridis, Stylianos I. Venieris, Mário Almeida, Ilias Leontiadis et al.MobiCom 2020 · 312 citations
- PatDNN: Achieving Real-Time DNN Execution on Mobile Devices with Pattern-based Weight PruningWei Niu, Xiaolong Ma, Sheng Lin, Shihao Wang et al.ASPLOS 2020 · 214 citations
- NEMO: enabling neural-enhanced video streaming on commodity mobile devicesHyunho Yeo, Chan Ju Chong, Youngmok Jung, Juncheol Ye et al.MobiCom 2020 · 118 citations
- Heimdall: mobile GPU coordination platform for augmented reality applicationsJuheon Yi, Youngki LeeMobiCom 2020 · 71 citations
Related papers
- HarDNet: A Low Memory Traffic NetworkPing Chao, Chao-Yang Kao, Yu-Shan Ruan, Chien-Hsiang Huang et al.ICCV 2019 · 303 citations
- Efficient Track AnythingYunyang Xiong, Chong Zhou, Xiaoyu Xiang, Lemeng Wu et al.ICCV 2025 · 5 citations
- Strata: Hierarchical Context Caching for Long Context Language Model ServingZhiqiang Xie, Ziyi Xu, Mark Zhao, Yuwei An et al.OSDI 2026 · 40 citations
- LouisKV: Efficient KV Cache Retrieval for Long Input-Output SequencesWenbo Wu, Qingyi Si, Xiurui Pan, Ye Wang et al.ICLR 2026 · 5 citations
- SmartCache: Context-aware Semantic Cache for Efficient Multi-turn LLM InferenceChengye Yu, Tianyu Wang, Zili Shao, Song JiangNeurIPS 2025 · 6 citations
