LMM-Det: Make Large Multimodal Models Excel in Object Detection
Jincheng Li, Chunyu Xie, Ji Ao, Dawei Leng, Yuhui Yin
摘要
Large multimodal models (LMMs) have garnered wide-spread attention and interest within the artificial intelligence research and industrial communities, owing to their remarkable capability in multimodal understanding, reasoning, and in-context learning, among others. While LMMs have demonstrated promising results in tackling multimodal tasks like image captioning, visual question answering, and visual grounding, the object detection capabilities of LMMs exhibit a significant gap compared to specialist detectors. To bridge the gap, we depart from the conventional methods of integrating heavy detectors with LMMs and propose LMM-Det, a simple yet effective approach that leverages a Large Multimodal Model for vanilla object Detection without relying on specialized detection modules. Specifically, we conduct a comprehensive exploratory analysis when a large multimodal model meets with object detection, revealing that the recall rate degrades significantly compared with specialist detection models. To mitigate this, we propose to increase the recall rate by introducing data distribution adjustment and inference optimization tailored for object detection. We re-organize the instruction conversations to enhance the object detection capabilities of large multimodal models. We claim that a large multimodal model possesses detection capability without any extra detection modules. Extensive experiments support our claim and show the effectiveness of the versatile LMM-Det. The datasets, models, and codes are available at https://github.com/360CVGroup/LMM-Det.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- WeDetect: Fast Open-Vocabulary Object Detection as RetrievalShenghao Fu, Yukun Su, Fengyun Rao, Jing Lyu 等CVPR 2026 · 被引用 11 次
- Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous DrivingJianhua Han, Meng Tian, Jiangtong Zhu, Fan He 等CVPR 2026 · 被引用 10 次
- VGent: Visual Grounding via Modular Design for Disentangling Reasoning and PredictionWeitai Kang, Jason Kuen, Mengwei Ren, Zijun Wei 等CVPR 2026 · 被引用 5 次
- RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided CaptioningJiahe Song, Chuang Wang, Bowen Jiang, Yinfan Wang 等CVPR 2026 · 被引用 3 次
- Rethinking Intermediate Representation for VLM-based Robot ManipulationWeiliang Tang, Jialin Gao, Jia-Hui Pan, Gang Wang 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- ROD-MLLM: Towards More Reliable Object Detection in Multimodal Large Language ModelsHeng Yin, Yuqiang Ren, Ke Yan, Shouhong Ding 等CVPR 2025
- Lumen: Unleashing Versatile Vision-Centric Capabilities of Large Multimodal ModelsYang Jiao, Shaoxiang Chen, Zequn Jie, Jingjing Chen 等NeurIPS 2024 · 被引用 27 次
- DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language ModelsYudong Zhang, Ruobing Xie, Xingwu Sun, Yiqing Huang 等ACM MM 2025 · 被引用 1 次
- How Can Objects Help Video-Language Understanding?Zitian Tang, Shijie Wang, Junho Cho, Jaewook Yoo 等ICCV 2025 · 被引用 8 次
- LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language ModelsShenghao Fu, Qize Yang, Qijie Mo, Junkai Yan 等CVPR 2025
