Hugging Visual Prompt and Segmentation Tokens: Consistency Learning for Fine-Grained Visual Understanding in MLLMs
jing yang, Sen Yang, Boqiang Duan, Ming Dai, Wei Zhang, Xiao Tan, Kunbin Chen, Wei He, Jingdong Wang, Hanli Wang
摘要
Recently, multimodal large language models (MLLMs) have achieved remarkable success in general multimodal tasks. Increasing attention has been given to leveraging MLLMs for fine-grained visual understanding, such as region-level captioning and pixel-level grounding.However, most existing approaches are task-specific, and some recent unified approaches attempt to handle both types simultaneously; they still fall short of deeply exploring the underlying associations across tasks. To bridge this gap, we propose a multimodal large language model designed to jointly support visual understanding through (FCLM). The central idea of this work is that pixel-level captioning and grounding are mutually beneficial and complementary tasks, each enhancing the other in achieving a fine-grained understanding of visual content.Specifically, FCLM analyzes the representation features -- visual prompt and segmentation tokens -- required for the two types of visual tasks, and achieves advanced reasoning and perception through a novel-designed consistency learning loss and a two-stage training framework. Moreover, we design a Hybrid Region Extractor to enhance the quality of visual prompt embeddings, thereby obtaining more semantically discriminative representations for detailed caption generation. Additionally, to verify the MLLM’s ability to localize accurate targets from detailed textual descriptions, we introduce a novel task called Detailed Localized Referring Expression Segmentation (DL-RES).We conduct extensive experiments on seven visual understanding tasks, demonstrating the strong performance and generalization ability of FCLM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper42
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Segment Everything Everywhere All at OnceXueyan Zou, Jianwei Yang, Hao Zhang, Feng Li 等NeurIPS 2023 · 被引用 889 次
- Ferret: Refer and Ground Anything Anywhere at Any GranularityHaoxuan You, Haotian Zhang, Zhe Gan, Xianzhi Du 等ICLR 2024 · 被引用 515 次
相关 Paper
- UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual ReasoningYe Liu, Zongyang Ma, Junfu Pu, Zhongang Qi 等NeurIPS 2025 · 被引用 39 次
- X-SAM: From Segment Anything to Any SegmentationHao Wang, Limeng Qiao, Zequn Jie, Zhijian Huang 等AAAI 2026 · 被引用 16 次
- Multi-Modal Instruction Tuned LLMs with Fine-Grained Visual PerceptionJunwen He, Yifan Wang, Lijun Wang, Huchuan Lu 等CVPR 2024
- Task-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual GroundingWenbo Chen, Zhen Xu, Ruotao Xu, Si Wu 等CVPR 2025
- Contrastive Localized Language-Image Pre-TrainingHong-You Chen, Zhengfeng Lai, Haotian Zhang, Xinze Wang 等ICML 2025
