Boosting Mobile CNN Inference through Semantic Memory
Yun Li, Chen Zhang, Shihao Han, Li Lyna Zhang, Baoqun Yin, Yunxin Liu, Mengwei Xu
摘要
Human brains are known to be capable of speeding up visual recognition of repeatedly presented objects through faster memory encoding and accessing procedures on activated neurons. For the first time, we borrow and distill such a capability into a semantic memory design, namely SMTM, to improve on-device CNN inference. SMTM employs a hierarchical memory architecture to leverage the long-tail distribution of objects of interest, and further incorporates several novel techniques to put it into effects: (1) it encodes high-dimensional feature maps into low-dimensional, semantic vectors for low-cost yet accurate cache and lookup; (2) it uses a novel metric in determining the exit timing considering different layers' inherent characteristics; (3) it adaptively adjusts the cache size and semantic vectors to fit the scene dynamics. SMTM is prototyped on commodity CNN engine and runs on both mobile CPU and GPU. Extensive experiments on large-scale datasets and models show that SMTM can significantly speed up the model inference over standard approach (up to 2×) and prior cache designs (up to 1.5x), with acceptable accuracy loss.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- LUT-NN: Empower Efficient Neural Network Inference with Centroid Learning and Table LookupXiaohu Tang, Yang Wang, Ting Cao, Li Lyna Zhang 等MobiCom 2023 · 被引用 29 次
- LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning AcceleratorGuoyu Li, Shengyu Ye, Chunyun Chen, Yang Wang 等HPCA 2025 · 被引用 7 次
- Accelerating End-Cloud Collaborative Inference via Near Bubble-Free Pipeline OptimizationLuyao Gao, Jianchun Liu, Hongli Xu, Sun Xu 等INFOCOM 2025 · 被引用 4 次
- Many Hands Make Light Work: Accelerating Edge Inference via Multi-Client Collaborative CachingWenyi Liang, Jianchun Liu, Hongli Xu, Chunming Qiao 等ICDE 2025
它引用的顶会 Paper6
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo 等ICCV 2019 · 被引用 633 次
- SPINN: synergistic progressive inference of neural networks over device and cloudStefanos Laskaridis, Stylianos I. Venieris, Mário Almeida, Ilias Leontiadis 等MobiCom 2020 · 被引用 312 次
- PatDNN: Achieving Real-Time DNN Execution on Mobile Devices with Pattern-based Weight PruningWei Niu, Xiaolong Ma, Sheng Lin, Shihao Wang 等ASPLOS 2020 · 被引用 214 次
- NEMO: enabling neural-enhanced video streaming on commodity mobile devicesHyunho Yeo, Chan Ju Chong, Youngmok Jung, Juncheol Ye 等MobiCom 2020 · 被引用 118 次
- Heimdall: mobile GPU coordination platform for augmented reality applicationsJuheon Yi, Youngki LeeMobiCom 2020 · 被引用 71 次
相关 Paper
- HarDNet: A Low Memory Traffic NetworkPing Chao, Chao-Yang Kao, Yu-Shan Ruan, Chien-Hsiang Huang 等ICCV 2019 · 被引用 303 次
- Efficient Track AnythingYunyang Xiong, Chong Zhou, Xiaoyu Xiang, Lemeng Wu 等ICCV 2025 · 被引用 5 次
- Strata: Hierarchical Context Caching for Long Context Language Model ServingZhiqiang Xie, Ziyi Xu, Mark Zhao, Yuwei An 等OSDI 2026 · 被引用 40 次
- LouisKV: Efficient KV Cache Retrieval for Long Input-Output SequencesWenbo Wu, Qingyi Si, Xiurui Pan, Ye Wang 等ICLR 2026 · 被引用 5 次
- SmartCache: Context-aware Semantic Cache for Efficient Multi-turn LLM InferenceChengye Yu, Tianyu Wang, Zili Shao, Song JiangNeurIPS 2025 · 被引用 6 次
