GroundingGPT: Language Enhanced Multi-modal Grounding Model
Zhaowei Li, Qi Xu, Dong Zhang, Hang Song, Yiqing Cai, Qi Qi, Ran Zhou, Junting Pan, Zefeng Li, Vu Tu, Zhida Huang, Tao Wang
摘要
Multi-modal large language models (MLLMs) have demonstrated remarkable performance across various tasks. However, these models often prioritize capturing global information and overlook the importance of perceiving local information. This limitation hinders their ability to effectively understand fine-grained details and handle grounding tasks that necessitate nuanced comprehension. Although some recent works have made strides in this, they have primarily focused on single-modality inputs. Therefore, we propose Grounding-GPT, an end-to-end language enhanced multimodal grounding model. It is designed to perform fine-grained grounding tasks for three modalities: image, video and audio. To enhance the model's performance, we adopt a coarse-to-fine training strategy, utilizing a threestage training approach to progressively enhance the model's semantic awareness and finegrained understanding capabilities. Additionally, we employ a diversified stage-specific dataset construction pipeline, developing a multi-modal, multi-granularity dataset tailored for training the model in different stages. Extensive experiments conducted on multiple multimodal benchmarks demonstrate that our model achieves impressive fine-grained understanding of multi-modal inputs on grounding tasks while maintaining or improving its global comprehension capabilities. Our code, model, and dataset are available at https://github. com/lzw-lzw/GroundingGPT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper41
- SpeechAlign: Aligning Speech Generation to Human PreferencesDong Zhang, Zhaowei Li, Shimin Li, Xin Zhang 等NeurIPS 2024 · 被引用 74 次
- SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal FusionMing Dai, Lingfeng Yang, Yihao Xu, Zhenhua Feng 等NeurIPS 2024 · 被引用 67 次
- OneThinker: All-in-one Reasoning Model for Image and VideoKaituo Feng, Manyuan Zhang, Hongyu Li, Kaixuan Fan 等CVPR 2026 · 被引用 55 次
- Artemis: Towards Referential Understanding in Complex VideosJihao Qiu, Yuan Zhang, Xi Tang, Lingxi Xie 等NeurIPS 2024 · 被引用 33 次
- TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMsYunheng Li, Jing Cheng, Shaoyong Jia, Hangyi Kuang 等NeurIPS 2025 · 被引用 18 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
相关 Paper
- Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal ModelsWei Wang, Zhaowei Li, Qi Xu, Linfeng Li 等EMNLP 2025 · 被引用 1 次
- Hugging Visual Prompt and Segmentation Tokens: Consistency Learning for Fine-Grained Visual Understanding in MLLMsjing yang, Sen Yang, Boqiang Duan, Ming Dai 等CVPR 2026
- On Pursuit of Designing Multi-modal Transformer for Video GroundingMeng Cao, Long Chen, Mike Zheng Shou, Can Zhang 等EMNLP 2021 · 被引用 63 次
- Enrich and Detect: Video Temporal Grounding With Multimodal LlmsShraman Pramanick, Effrosyni Mavroudi, Yale Song, Rama Chellappa 等ICCV 2025 · 被引用 4 次
- Grounding-IQA: Grounding Multimodal Language Model for Image Quality AssessmentZheng Chen, Xun Zhang, Wenbo Li, Renjing Pei 等ICLR 2026 · 被引用 12 次
