Visual Perception by Large Language Model's Weights
Feipeng Ma, Hongwei Xue, Yizhou Zhou, Guangting Wang, Fengyun Rao, Shilin Yan, Yueyi Zhang, Siying Wu, Mike Zheng Shou, Xiaoyan Sun
Abstract
Existing Multimodal Large Language Models (MLLMs) follow the paradigm that perceives visual information by aligning visual features with the input space of Large Language Models (LLMs), and concatenating visual tokens with text tokens to form a unified sequence input for LLMs. These methods demonstrate promising results on various vision-language tasks but are limited by the high computational effort due to the extended input sequence resulting from the involvement of visual tokens. In this paper, instead of input space alignment, we propose a novel parameter space alignment paradigm that represents visual information as model weights. For each input image, we use a vision encoder to extract visual features, convert features into perceptual weights, and merge the perceptual weights with LLM's weights. In this way, the input of LLM does not require visual tokens, which reduces the length of the input sequence and greatly improves efficiency. Following this paradigm, we propose VLoRA with the perceptual weights generator. The perceptual weights generator is designed to convert visual features to perceptual weights with low-rank property, exhibiting a form similar to LoRA. The experimental results show that our VLoRA achieves comparable performance on various benchmarks for MLLMs, while significantly reducing the computational costs for both training and inference. The code and models will be made open-source.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 36e62081-1d84-44b9-9be3-22fd059cfa56Cited by top-tier papers8
- CoRA: Collaborative Information Perception by Large Language Model's Weights for RecommendationYuting Liu, Jinghao Zhang, Yizhou Dang, Yuliang Liang et al.AAAI 2025 · 15 citations
- HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language ModelsZelin Peng, Zhengqin Xu, Qingyang Liu, Xiaokang Yang et al.NeurIPS 2025 · 5 citations
- IF-Prune: Information-Flow Guided Token Pruning for Efficient Vision-Language ModelsGuohao Sun, Yufei Wang, Sizhuo Ma, Yuege Xie et al.CVPR 2026
- CP-CLIP: Customized Parameter Generation for Open-vocabulary Semantic SegmentationZelin Peng, Zhengqin Xu, Feilong Tang, Wei ShenAAAI 2026
- ViPE: Visual Perception in Parameter Space for Efficient Video-Language UnderstandingShichen Lu, Tongtian Yue, Longteng Guo, Handong Li et al.EMNLP 2025
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- ROSE: Rotate Your Large Language Model to SeeTongtian Yue, Xuange Gao, Longteng Guo, Zijia Zhao et al.CVPR 2026
- LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-SteeringJinhe Bi, Yujun Wang, Haokun Chen, Xun Xiao et al.ACL 2025
- Empower Vision Applications with LoRA LMMLiang Mi, Weijun Wang, Wenming Tu, Qingfeng He et al.EuroSys 2025 · 2 citations
- AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document UnderstandingAhmed Masry, Juan A. Rodríguez, Tianyu Zhang, Suyuchen Wang et al.NeurIPS 2025 · 7 citations
- Deep Pre-Alignment for VLMsTianyu Yu, Kechen Fang, Zihao Wan, Kaidong Zhang et al.ICML 2026
