GroundingGPT: Language Enhanced Multi-modal Grounding Model
Zhaowei Li, Qi Xu, Dong Zhang, Hang Song, Yiqing Cai, Qi Qi, Ran Zhou, Junting Pan, Zefeng Li, Vu Tu, Zhida Huang, Tao Wang
Abstract
Multi-modal large language models (MLLMs) have demonstrated remarkable performance across various tasks. However, these models often prioritize capturing global information and overlook the importance of perceiving local information. This limitation hinders their ability to effectively understand fine-grained details and handle grounding tasks that necessitate nuanced comprehension. Although some recent works have made strides in this, they have primarily focused on single-modality inputs. Therefore, we propose Grounding-GPT, an end-to-end language enhanced multimodal grounding model. It is designed to perform fine-grained grounding tasks for three modalities: image, video and audio. To enhance the model's performance, we adopt a coarse-to-fine training strategy, utilizing a threestage training approach to progressively enhance the model's semantic awareness and finegrained understanding capabilities. Additionally, we employ a diversified stage-specific dataset construction pipeline, developing a multi-modal, multi-granularity dataset tailored for training the model in different stages. Extensive experiments conducted on multiple multimodal benchmarks demonstrate that our model achieves impressive fine-grained understanding of multi-modal inputs on grounding tasks while maintaining or improving its global comprehension capabilities. Our code, model, and dataset are available at https://github. com/lzw-lzw/GroundingGPT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers41
- SpeechAlign: Aligning Speech Generation to Human PreferencesDong Zhang, Zhaowei Li, Shimin Li, Xin Zhang et al.NeurIPS 2024 · 74 citations
- SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal FusionMing Dai, Lingfeng Yang, Yihao Xu, Zhenhua Feng et al.NeurIPS 2024 · 67 citations
- OneThinker: All-in-one Reasoning Model for Image and VideoKaituo Feng, Manyuan Zhang, Hongyu Li, Kaixuan Fan et al.CVPR 2026 · 55 citations
- Artemis: Towards Referential Understanding in Complex VideosJihao Qiu, Yuan Zhang, Xi Tang, Lingxi Xie et al.NeurIPS 2024 · 33 citations
- TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMsYunheng Li, Jing Cheng, Shaoyong Jia, Hangyi Kuang et al.NeurIPS 2025 · 18 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
Related papers
- Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal ModelsWei Wang, Zhaowei Li, Qi Xu, Linfeng Li et al.EMNLP 2025 · 1 citation
- Hugging Visual Prompt and Segmentation Tokens: Consistency Learning for Fine-Grained Visual Understanding in MLLMsjing yang, Sen Yang, Boqiang Duan, Ming Dai et al.CVPR 2026
- On Pursuit of Designing Multi-modal Transformer for Video GroundingMeng Cao, Long Chen, Mike Zheng Shou, Can Zhang et al.EMNLP 2021 · 63 citations
- Enrich and Detect: Video Temporal Grounding With Multimodal LlmsShraman Pramanick, Effrosyni Mavroudi, Yale Song, Rama Chellappa et al.ICCV 2025 · 4 citations
- Grounding-IQA: Grounding Multimodal Language Model for Image Quality AssessmentZheng Chen, Xun Zhang, Wenbo Li, Renjing Pei et al.ICLR 2026 · 12 citations
