RegionGPT: Towards Region Understanding Vision Language Model
Qiushan Guo, Shalini De Mello, Hongxu Yin, Wonmin Byeon, Ka Chun Cheung, Yizhou Yu, Ping Luo, Sifei Liu
2024Year
45Top-tier citations
Abstract
Figure 1 . We introduce RegionGPT that enables complex region-level captioning, reasoning, classification, and expression comprehension capabilities for the multimodal large language model. Users can input regions of interest of any shape, utilizing ⟨region⟩ as a placeholder within the instruction at any position. Such placeholders are subsequently replaced with semantic region-level embeddings that are fed into the language decoder. Best viewed in color.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers45
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo et al.NeurIPS 2024 · 412 citations
- DICEPTION: A Generalist Diffusion Model for Visual Perceptual TasksCanyu Zhao, Yanlong Sun, Mingyu Liu, Huanyi Zheng et al.NeurIPS 2025 · 45 citations
- Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMsYifan Shen, Yuanzhe Liu, Jingyuan Zhu, Xu Cao et al.NeurIPS 2025 · 41 citations
- 3D Aware Region Prompted Vision Language ModelAn-Chieh Cheng, Yang Fu, Yukang Chen, Zhijian Liu et al.ICLR 2026 · 30 citations
- PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action ModelWenqi Liang, Gan Sun, Yao He, Jiahua Dong et al.ICLR 2026 · 20 citations
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMsHaochen Wang, Yuhao Wang, Tao Zhang, Yikang Zhou et al.ICLR 2026 · 18 citations
- What does CLIP know about a red circle? Visual prompt engineering for VLMsAleksandar Shtedritski, Christian Rupprecht, Andrea VedaldiICCV 2023 · 262 citations
- NExT-Chat: An LMM for Chat, Detection and SegmentationAo Zhang, Yuan Yao, Wei Ji, Zhiyuan Liu et al.ICML 2024 · 85 citations
- MedSIGHT: Towards Grounded Visual Comprehension in Medical Large Vision-Language ModelsAofei Chang, Le Huang, Alex Boyd, parminder bhatia et al.ICML 2026
- KptLLM: Unveiling the Power of Large Language Model for Keypoint ComprehensionJie Yang, Wang Zeng, Sheng Jin, Lumin Xu et al.NeurIPS 2024 · 9 citations
