mVLM: A Vision Language Model for mNPUs
Zijie Chen, Guiyun Fan, Zhaoxing Yang, Rong Ding, Haiming Jin
Abstract
The proliferation of low-power intelligent processors with integrated Neural Processing Units (NPUs), called µNPUs, has created new opportunities for on-device generative AI, benefitting end devices like smart wearables and small robots. However, deploying Vision-Language Models (VLMs) on µNPUs is severely hindered by stringent memory constraints and limited operator support. To bridge this critical gap, we propose µVLM, the first lightweightoriented VLM architecture designed for µNPUs. It is comprised of our proposed OverMod encoder and AttSSM decoder. OverMod is a lightweight dynamic convolutional network inspired by biomimetic vision, incorporating our novel Global Spatial Modulation mechanism to enable adaptive, high-fidelity feature extraction using only NPUfriendly operators. AttSSM leverages a highly efficient State Space Model (SSM) core, augmented with multi-scale feature fusion and Global Context Dynamic Modulation mechanism, to perform robust sequential modeling. Furthermore, we introduce a coordinated full-parameter quantization strategy that preserves precision across the encoderdecoder boundary, alongside hand-optimized operators for unsupported modules like SSMs. µVLM achieves a competitive CIDEr score of 117.8 on the COCO Karpathy test split and, for the first time, demonstrates the feasibility of millisecond-level VLM inference on a µNPU platform.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cca7a2ff-c6f2-4d6a-8a30-429c5ba11788Builds on19
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve et al.ICCV 2021 · 1,279 citations
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 992 citations
- EfficientFormer: Vision Transformers at MobileNet SpeedYanyu Li, Geng Yuan, Yang Wen, Ju Hu et al.NeurIPS 2022 · 742 citations
Related papers
- AQuant: Repurposing CODEC for VLM Acceleration via Adaptive QuantizationZhuoran Song, Chunyu Qi, Jian Weng, Xiaoyao Liang et al.ISCA 2026
- SPEED-Q: Staged Processing with Enhanced Distillation Towards Efficient Low-Bit On-Device VLM QuantizationTianyu Guo, Shanwei Zhao, Shiai Zhu, Chenguang MaAAAI 2026
- DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot ExecutionYang Yue, Yulin Wang, Bingyi Kang, Yizeng Han et al.NeurIPS 2024 · 153 citations
- BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile DevicesXudong Lu, Yinghao Chen, Cheng Chen, Hui Tan et al.CVPR 2025
- QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language ModelsYutong Wang, Haiyu Wang, Sai Qian ZhangNeurIPS 2025 · 16 citations
