mVLM: A Vision Language Model for mNPUs
Zijie Chen, Guiyun Fan, Zhaoxing Yang, Rong Ding, Haiming Jin
摘要
The proliferation of low-power intelligent processors with integrated Neural Processing Units (NPUs), called µNPUs, has created new opportunities for on-device generative AI, benefitting end devices like smart wearables and small robots. However, deploying Vision-Language Models (VLMs) on µNPUs is severely hindered by stringent memory constraints and limited operator support. To bridge this critical gap, we propose µVLM, the first lightweightoriented VLM architecture designed for µNPUs. It is comprised of our proposed OverMod encoder and AttSSM decoder. OverMod is a lightweight dynamic convolutional network inspired by biomimetic vision, incorporating our novel Global Spatial Modulation mechanism to enable adaptive, high-fidelity feature extraction using only NPUfriendly operators. AttSSM leverages a highly efficient State Space Model (SSM) core, augmented with multi-scale feature fusion and Global Context Dynamic Modulation mechanism, to perform robust sequential modeling. Furthermore, we introduce a coordinated full-parameter quantization strategy that preserves precision across the encoderdecoder boundary, alongside hand-optimized operators for unsupported modules like SSMs. µVLM achieves a competitive CIDEr score of 117.8 on the COCO Karpathy test split and, for the first time, demonstrates the feasibility of millisecond-level VLM inference on a µNPU platform.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,279 次
- Attention on Attention for Image CaptioningLun Huang, Wenmin Wang, Jie Chen, Xiaoyong WeiICCV 2019 · 被引用 992 次
- EfficientFormer: Vision Transformers at MobileNet SpeedYanyu Li, Geng Yuan, Yang Wen, Ju Hu 等NeurIPS 2022 · 被引用 742 次
相关 Paper
- AQuant: Repurposing CODEC for VLM Acceleration via Adaptive QuantizationZhuoran Song, Chunyu Qi, Jian Weng, Xiaoyao Liang 等ISCA 2026
- SPEED-Q: Staged Processing with Enhanced Distillation Towards Efficient Low-Bit On-Device VLM QuantizationTianyu Guo, Shanwei Zhao, Shiai Zhu, Chenguang MaAAAI 2026
- DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot ExecutionYang Yue, Yulin Wang, Bingyi Kang, Yizeng Han 等NeurIPS 2024 · 被引用 153 次
- BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile DevicesXudong Lu, Yinghao Chen, Cheng Chen, Hui Tan 等CVPR 2025
- QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language ModelsYutong Wang, Haiyu Wang, Sai Qian ZhangNeurIPS 2025 · 被引用 16 次
