BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices
Xudong Lu, Yinghao Chen, Cheng Chen, Hui Tan, Boheng Chen, Yina Xie, Rui Hu, Guanxin Tan, Renshou Wu, Yan Hu, Yi Zeng, Lei Wu
Abstract
The emergence and growing popularity of multimodal large language models (MLLMs) have significant potential to enhance various aspects of daily life, from improving communication to facilitating learning and problem-solving. Mobile phones, as essential daily companions, represent the most effective and accessible deployment platform for MLLMs, enabling seamless integration into everyday tasks. However, deploying MLLMs on mobile phones presents challenges due to limitations in memory size and computational capability, making it difficult to achieve smooth and real-time processing without extensive optimization. In this paper, we present BlueLM-V-3B, an algorithm and system co-design approach specifically tailored for the efficient deployment of MLLMs on mobile platforms. To be specific, we redesign the dynamic resolution scheme adopted by mainstream MLLMs and implement system optimization for hardware-aware deployment to optimize model inference on mobile phones. BlueLM-V-3B boasts the following key highlights: (1) Small Size: BlueLM-V-3B features a language model with 2.7B parameters and a vision encoder with 400M parameters. (2) Fast Speed: BlueLM-V-3B achieves a generation speed of 24.4 token/s on the MediaTek Dimensity 9300 processor with 4-bit LLM weight quantization. (3) Strong Performance: BlueLM-V-3B has attained the highest average score of 66.1 on the OpenCompass benchmark among models with ≤ 4B parameters and surpassed a series of models with much larger parameter sizes (e.g., MiniCPM-V-2.6, InternVL2-8B).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f86c2a60-8d07-43e8-81cd-34c5df4fee67Cited by top-tier papers8
- Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language ModelsXuyang Liu, Ziming Wang, Junjie Chen, Yuhang Han et al.AAAI 2026 · 25 citations
- PhoStream: Benchmarking Real-World Streaming for Omnimodal Assistants in Mobile ScenariosXudong LU, Guan Huankang, Yang Bo, Jinpeng Chen et al.ICML 2026 · 9 citations
- Scaling LLM Test-Time Compute with Mobile NPU on SmartphonesZixu Hao, Jianyu Wei, Tuowei Wang, Minxing Huang et al.EuroSys 2026 · 2 citations
- COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile DevicesYilong Zhao, Fangxin Liu, Onur Mutlu, Mingyu Gao et al.ISCA 2026 · 1 citation
- GenieBlue: Integrating Both Linguistic and Multimodal Capabilities for Large Language Models on Mobile DevicesXudong Lu, Yinghao Chen, Renshou Wu, Haohao Gao et al.ICCV 2025 · 1 citation
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang et al.ICML 2024 · 1,191 citations
Related papers
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training RecipeTianyu Yu, Zefan Wang, Chongyi Wang, Fuwei Huang et al.CVPR 2026 · 179 citations
- MQuant: Unleashing the Inference Potential of Multimodal Large Language Models via Static QuantizationJiangyong Yu, Sifan Zhou, Dawei Yang, Shuoyu Li et al.ACM MM 2025 · 11 citations
- Efficient Multimodal Large Language Model via Dynamic KV Cache QuantizationJiahao Fan, Chien-Ming ChenAAAI 2026
- MBQ: Modality-Balanced Quantization for Large Vision-Language ModelsShiyao Li, Yingchun Hu, Xuefei Ning, Xihui Liu et al.CVPR 2025
- BitMoD: Bit-serial Mixture-of-Datatype LLM AccelerationYuzong Chen, Ahmed F. AbouElhamayed, Xilai Dai, Yang Wang et al.HPCA 2025 · 23 citations
