DetGPT: Detect What You Need via Reasoning
Renjie Pi, Jiahui Gao, Shizhe Diao, Rui Pan, Hanze Dong, Jipeng Zhang, Lewei Yao, Jianhua Han, Hang Xu, Lingpeng Kong, Tong Zhang
Abstract
Recently, vision-language models (VLMs) such as GPT4, LLAVA, and MiniGPT4 have witnessed remarkable breakthroughs, which are great at generating image descriptions and visual question answering. However, it is difficult to apply them to an embodied agent for completing real-world tasks, such as grasping, since they can not localize the object of interest. In this paper, we introduce a new task termed reasoning-based object detection, which aims at localizing the objects of interest in the visual scene based on any human instructs. Our proposed method, called DetGPT, leverages instruction-tuned VLMs to perform reasoning and find the object of interest, followed by an open-vocabulary object detector to localize these objects. DetGPT can automatically locate the object of interest based on the user's expressed desires, even if the object is not explicitly mentioned. This ability makes our system potentially applicable across a wide range of fields, from robotics to autonomous driving. To facilitate research in the proposed reasoningbased object detection, we curate and opensource a benchmark named RD-Bench for instruction tuning and evaluation. Overall, our proposed task and DetGPT demonstrate the potential for more sophisticated and intuitive interactions between humans and machines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7669ae4-a72d-48e1-a7d3-c9a25062b45eCited by top-tier papers29
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial IntelligenceDiankun Wu, Fangfu Liu, Yi-Hsin Hung, Yueqi DuanNeurIPS 2025 · 245 citations
- One Token to Seg Them All: Language Instructed Reasoning Segmentation in VideosZechen Bai, Tong He, Haiyang Mei, Pichao Wang et al.NeurIPS 2024 · 147 citations
- Multi-Object Hallucination in Vision Language ModelsXuweiyi Chen, Ziqiao Ma, Xuejun Zhang, Sihan Xu et al.NeurIPS 2024 · 77 citations
- Universal Video Temporal Grounding with Generative Multi-modal Large Language ModelsZeqian Li, Shangzhe Di, Zhonghua Zhai, Weilin Huang et al.NeurIPS 2025 · 30 citations
- Chain of Visual Perception: Harnessing Multimodal Large Language Models for Zero-shot Camouflaged Object DetectionLv Tang, Peng-Tao Jiang, Zhihao Shen, Hao Zhang et al.ACM MM 2024 · 29 citations
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi et al.ICCV 2019 · 3,348 citations
Related papers
- ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance ConstraintsPei-An Chen, Yong-Ching Liang, Jia-Fong Yeh, Hung-Ting Su et al.ACL 2026 · 1 citation
- NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language ModelsGengze Zhou, Yicong Hong, Qi WuAAAI 2024 · 361 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Ins-DetCLIP: Aligning Detection Model to Follow Human-Language InstructionRenjie Pi, Lewei Yao, Jianhua Han, Xiaodan Liang et al.ICLR 2024 · 5 citations
- Vision-Language-Action Instruction Tuning: From Understanding to ManipulationShuai Yang, Hao Li, Bin Wang, Yilun Chen et al.ICLR 2026 · 50 citations
