A Comprehensive Overhaul of Multimodal Assistant with Small Language Models
Minjie Zhu, Yichen Zhu, Ning Liu, Xin Liu, Zhiyuan Xu, Chaomin Shen, Yaxin Peng
摘要
Multimodal Large Language Models (MLLMs) have showcased impressive skills in tasks related to visual understanding and reasoning. Yet, their widespread application faces obstacles due to the high computational demands during both the training and inference phases, restricting their use to a limited audience within the research and user communities. In this paper, we investigate the design aspects of Multimodal Small Language Models (MSLMs) and propose an efficient multimodal assistant named Mipha, which is designed to create synergy among various aspects: visual representation, language models, and optimization strategies. We show that without increasing the volume of training data, our Mipha-3B outperforms the state-of-the-art large MLLMs, especially LLaVA-1.5-13B, on multiple benchmarks. Through detailed discussion, we provide insights and guidelines for developing strong MSLMs that rival the capabilities of MLLMs. Our code is available at https://github.com/zhuyiche/llava-phi .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing ImagesZepeng Xin, Kaiyu Li, Luodi Chen, Wanchen Li 等CVPR 2026 · 被引用 14 次
- Instructseg: Unifying Instructed Visual Segmentation with Multi-Modal Large Language ModelsCong Wei, Yujie Zhong, Haoxian Tan, Yingsen Zeng 等ICCV 2025 · 被引用 8 次
- Adadrive: Self-Adaptive Slow-Fast System for Language-Grounded Autonomous DrivingRuifei Zhang, Junlin Xie, Wei Zhang, Weikai Chen 等ICCV 2025 · 被引用 3 次
- Any2Policy: Learning Visuomotor Policy with Any-ModalityYichen Zhu, Zhicai Ou, Feifei Feng, Jian TangNeurIPS 2024 · 被引用 3 次
- FlashSloth : Lightning Multimodal Large Language Models via Embedded Visual CompressionBo Tong, Bokai Lai, Yiyi Zhou, Gen Luo 等CVPR 2025
它引用的顶会 Paper18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
相关 Paper
- Improved Baselines with Visual Instruction TuningHaotian Liu, Chunyuan Li, Yuheng Li, Yong Jae LeeCVPR 2024
- Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language ModelsGen Luo, Yiyi Zhou, Yuxin Zhang, Xiawu Zheng 等ICLR 2025
- Eve: Efficient Multimodal Vision Language Models with Elastic Visual ExpertsMiao Rang, Zhenni Bi, Chuanjian Liu, Yehui Tang 等AAAI 2025 · 被引用 16 次
- Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient InferenceHan Zhao, Min Zhang, Wei Zhao, Pengxiang Ding 等AAAI 2025 · 被引用 125 次
- LLaVA-KD: A Framework of Distilling Multimodal Large Language ModelsYuxuan Cai, Jiangning Zhang, Haoyang He, Xinwei He 等ICCV 2025 · 被引用 9 次
