Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
Gen Luo, Yiyi Zhou, Yuxin Zhang, Xiawu Zheng, Xiaoshuai Sun, Rongrong Ji
Abstract
Despite remarkable progress, existing multimodal large language models (MLLMs) are still inferior in granular visual recognition. Contrary to previous works, we study this problem from the perspective of image resolution, and reveal that a combination of low-and high-resolution visual features can effectively mitigate this shortcoming. Based on this observation, we propose a novel and efficient method for MLLMs, termed Mixture-of-Resolution Adaptation (MRA). In particular, MRA adopts two visual pathways for images with different resolutions, where highresolution visual information is embedded into the low-resolution pathway via the novel mixture-ofresolution adapters (MR-Adapters). This design also greatly reduces the input sequence length of MLLMs. To validate MRA, we apply it to a recent MLLM called LLaVA, and term the new model LLaVA-HR. We conduct extensive experiments on 11 vision-language (VL) tasks, which show that LLaVA-HR outperforms existing MLLMs on 8 VL tasks, e.g., +9.4% on TextVQA. More importantly, both training and inference of LLaVA-HR remain efficient with MRA, e.g., 20 training hours and 3× inference speed than LLaVA-1.5. Source codes are released at: https:// github.com/luogen1996/LLaVA-HR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers60
- Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMsQizhe Zhang, Mengzhen Liu, Lichen Li, Ming Lu et al.NeurIPS 2025 · 104 citations
- Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language ModelsWeihao Ye, Qiong Wu, Wenhao Lin, Yiyi ZhouAAAI 2025 · 99 citations
- ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language ModelsMingrui Wu, Xinyue Cai, Jiayi Ji, Jiale Li et al.NeurIPS 2024 · 50 citations
- Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context RetentionXin Zou, Di Lu, Yizhou Wang, Yibo Yan et al.NeurIPS 2025 · 49 citations
- Meteor: Mamba-based Traversal of Rationale for Large Language and Vision ModelsByung-Kwan Lee, Chae Won Kim, Beomchan Park, Yong Man RoNeurIPS 2024 · 37 citations
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
Related papers
- Cheap and Quick: Efficient Vision-Language Instruction Tuning for Large Language ModelsGen Luo, Yiyi Zhou, Tianhe Ren, Shengxin Chen et al.NeurIPS 2023 · 157 citations
- Task-Aware Resolution Optimization for Visual Large Language ModelsWeiqing Luo, Zhen Tan, Yifan Li, Xinyu Zhao et al.EMNLP 2025 · 1 citation
- Cross-modal Information Flow in Multimodal Large Language ModelsZhi Zhang, Srishti Yadav, Fengze Han, Ekaterina ShutovaCVPR 2025
- A Comprehensive Overhaul of Multimodal Assistant with Small Language ModelsMinjie Zhu, Yichen Zhu, Ning Liu, Xin Liu et al.AAAI 2025 · 30 citations
- Visual Perception by Large Language Model's WeightsFeipeng Ma, Hongwei Xue, Yizhou Zhou, Guangting Wang et al.NeurIPS 2024 · 24 citations
