Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal Models
Yang Jiao, Shaoxiang Chen, Zequn Jie, Jingjing Chen, Lin Ma, Yu-Gang Jiang
Abstract
Large Multimodal Model (LMM) is a hot research topic in the computer vision area and has also demonstrated remarkable potential across multiple disciplinary fields. A recent trend is to further extend and enhance the perception capabilities of LMMs. The current methods follow the paradigm of adapting the visual task outputs to the format of the language model, which is the main component of a LMM. This adaptation leads to convenient development of such LMMs with minimal modifications, however, it overlooks the intrinsic characteristics of diverse visual tasks and hinders the learning of perception capabilities. To address this issue, we propose a novel LMM architecture named Lumen, a Large multimodal model with versatile vision-centric capability enhancement. We decouple the LMM's learning of perception capabilities into task-agnostic and task-specific stages. Lumen first promotes fine-grained vision-language concept alignment, which is the fundamental capability for various visual tasks. Thus the output of the task-agnostic stage is a shared representation for all the tasks we address in this paper. Then the task-specific decoding is carried out by flexibly routing the shared representation to lightweight task decoders with negligible training efforts. Comprehensive experimental results on a series of vision-centric and VQA benchmarks indicate that our Lumen model not only achieves or surpasses the performance of existing LMM-based approaches in a range of vision-centric tasks while maintaining general visual understanding and instruction following capabilities. The code will be released at https://github.com/SxJyJay/Lumen.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 20a471f7-3536-40b9-911f-2085ad34465cCited by top-tier papers5
- CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept MatchingDongzhi Jiang, Guanglu Song, Xiaoshi Wu, Renrui Zhang et al.NeurIPS 2024 · 75 citations
- Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive ReasoningYian Li, Wentao Tian, Yang Jiao, Tianwen Qian et al.ACM MM 2025 · 16 citations
- SIFThinker: Spatially-Aware Image Focus for Visual ReasoningZhangquan Chen, Ruihui Zhao, Chuwei Luo, Mingze Sun et al.AAAI 2026 · 12 citations
- Semi-Supervised Clustering Framework for Fine-grained Scene Graph GenerationJiarui Yang, Chuan Wang, Jun Zhang, Shuyi Wu et al.AAAI 2025 · 2 citations
- Identity-Aware Vision-Language Model for Explainable Face Forgery DetectionJunhao Xu, Jingjing Chen, Yang Jiao, Jiacheng Zhang et al.AAAI 2026 · 1 citation
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- L-Man: A Large Multi-modal Model Unifying Human-centric TasksJialong Zuo, Ying Nie, Tianyu Guo, Huaxin Zhang et al.AAAI 2025 · 1 citation
- VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language TasksJiannan Wu, Muyan Zhong, Sen Xing, Zeqiang Lai et al.NeurIPS 2024 · 179 citations
- LMM-Det: Make Large Multimodal Models Excel in Object DetectionJincheng Li, Chunyu Xie, Ji Ao, Dawei Leng et al.ICCV 2025 · 2 citations
- Advancing Visual Large Language Model for Multi-Granular Versatile PerceptionWentao Xiang, Haoxian Tan, Yujie Zhong, Cong Wei et al.ICCV 2025 · 1 citation
- DenseMLLM: Standard Multimodal LLMs for Dense PredictionYi Li, Hongze Shen, Lexiang Tang, Xin Li et al.ICML 2026
