Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models
Wei Wang, Zhaowei Li, Qi Xu, Linfeng Li, Yiqing Cai, Botian Jiang, Hang Song, Xingcan Hu, Pengyu Wang, Li Xiao
摘要
Multi-modal large language models (MLLMs) have achieved remarkable success in finegrained visual understanding across a range of tasks. However, they often encounter significant challenges due to inadequate alignment for fine-grained knowledge, which restricts their ability to accurately capture local details and attain a comprehensive global perception. While recent advancements have focused on aligning object expressions with grounding information, they typically lack explicit integration of object images, which contain affluent information beyond mere texts or coordinates. To bridge this gap, we introduce a novel fine-grained visual knowledge alignment method that effectively aligns and integrates multi-scale knowledge of objects, including texts, coordinates, and images. This innovative method is underpinned by our multi-scale fine-grained enhancement data synthesis pipeline, which provides over 300K essential training data to enhance alignment and improve overall performance. Furthermore, we present TinyGroundingGPT, a series of compact models optimized for highlevel alignments. With a scale of approximately 3B parameters, TinyGroundingGPT achieves outstanding results in grounding tasks while delivering performance comparable to larger MLLMs in complex visual scenarios. The data and code will be released in https://github. com/wwangweii/TinyGroundingGPT.git .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image EditingWei Chow, Linfeng Li, Lingdong Kong, Zefeng Li 等CVPR 2026 · 被引用 14 次
- IAG: Input-aware Backdoor Attack on VLM-based Visual GroundingJunxian Li, Beining Xu, Simin Chen, Jiatong Li 等CVPR 2026 · 被引用 13 次
- Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language ModelsJiachen Jiang, Jinxin Zhou, Bo Peng, Xia Ning 等NeurIPS 2025 · 被引用 8 次
- MMKU-Bench: A Multimodal Update Benchmark for Diverse Visual KnowledgeBaochen Fu, Yuntao Du, Cheng Chang, Baihao Jin 等ICML 2026 · 被引用 7 次
- MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition QueryWei Chow, Yuan Gao, Linfeng Li, Xian Wang 等NeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
- CogVLM: Visual Expert for Pretrained Language ModelsWeihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong 等NeurIPS 2024 · 被引用 858 次
相关 Paper
- GroundingGPT: Language Enhanced Multi-modal Grounding ModelZhaowei Li, Qi Xu, Dong Zhang, Hang Song 等ACL 2024 · 被引用 29 次
- PostAlign: Multimodal Grounding as a Corrective Lens for MLLMsYixuan Wu, Yang Zhang, Jian Wu, Philip Torr 等ICLR 2026 · 被引用 5 次
- Hugging Visual Prompt and Segmentation Tokens: Consistency Learning for Fine-Grained Visual Understanding in MLLMsjing yang, Sen Yang, Boqiang Duan, Ming Dai 等CVPR 2026
- Connecting the Dots: Training-Free Visual Grounding via Agentic ReasoningLiqin Luo, Guangyao Chen, Xiawu Zheng, Yongxing Dai 等AAAI 2026
- SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding CapabilityJiankang Wang, Zhihan Zhang, Zhihang Liu, Yang Li 等AAAI 2026 · 被引用 20 次
