ROD-MLLM: Towards More Reliable Object Detection in Multimodal Large Language Models
Heng Yin, Yuqiang Ren, Ke Yan, Shouhong Ding, Yongtao Hao
摘要
Multimodal large language models (MLLMs) have demonstrated strong language understanding and generation capabilities, excelling in visual tasks like referring and grounding. However, due to task type limitations and dataset scarcity, existing MLLMs only ground objects present in images and cannot reject non-existent objects effectively, resulting in unreliable predictions. In this paper, we introduce ROD-MLLM, a novel MLLM for Reliable Object Detection using free-form language. We propose a query-based localization mechanism to extract low-level object features. By aligning global and object-level visual information with text space, we leverage the large language model (LLM) for high-level comprehension and final localization decisions, overcoming the language understanding limitations of normal detectors. To enhance languagebased object detection, we design an automated data annotation pipeline and construct the dataset ROD. This pipeline uses the referring capabilities of existing MLLMs and chain-of-thought techniques to generate diverse expressions corresponding to zero or multiple objects, addressing the shortage of training data. Experiments across various tasks, including referring, grounding, and language-based object detection, show that ROD-MLLM achieves state-ofthe-art performance among MLLMs. Notably, in languagebased object detection, our model achieves +13.7 AP improvement on D 3 benchmark over existing MLLMs and surpasses most specialized detection models, especially in scenarios requiring complex language understanding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- SAM 3: Segment Anything with ConceptsNicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath 等ICLR 2026 · 被引用 1,103 次
- WeDetect: Fast Open-Vocabulary Object Detection as RetrievalShenghao Fu, Yukun Su, Fengyun Rao, Jing Lyu 等CVPR 2026 · 被引用 11 次
- See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying TogglesZongru Wu, Rui Mao, Zhiyuan Tian, Pengzhou Cheng 等CVPR 2026 · 被引用 3 次
- ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason ChunkingLihong Wang, Liangqi Li, Weiwei Feng, Jiamin Wu 等CVPR 2026 · 被引用 1 次
- Fast SceneScript: Fast and Accurate Language‑Based 3D Scene Understanding via Multi‑Token PredictionRuihong Yin, Xuepeng Shi, Oleksandr Bailo, Marco Manfredi 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
相关 Paper
- Connecting the Dots: Training-Free Visual Grounding via Agentic ReasoningLiqin Luo, Guangyao Chen, Xiawu Zheng, Yongxing Dai 等AAAI 2026
- RealVG: Unleashing MLLMs for Training-Free Spatio-Temporal Video Grounding in the WildHongchen Wei, Zhenzhong ChenACM MM 2025 · 被引用 1 次
- GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional EvaluationRang Li, Lei Li, Shuhuai Ren, Hao Tian 等CVPR 2026 · 被引用 10 次
- Spatial Preference Rewarding for MLLMs Spatial UnderstandingHan Qiu, Peng Gao, Lewei Lu, Xiaoqin Zhang 等ICCV 2025 · 被引用 3 次
- InstructDET: Diversifying Referring Object Detection with Generalized InstructionsRonghao Dang, Jiangyan Feng, Haodong Zhang, Chongjian Ge 等ICLR 2024 · 被引用 16 次
