Chain of Visual Perception: Harnessing Multimodal Large Language Models for Zero-shot Camouflaged Object Detection
Lv Tang, Peng-Tao Jiang, Zhihao Shen, Hao Zhang, Jinwei Chen, Bo Li
Abstract
In this paper, we introduce a novel multimodal camo-perceptive framework (MMCPF) aimed at handling zero-shot Camouflaged Object Detection (COD) by leveraging the powerful capabilities of Multimodal Large Language Models (MLLMs). Recognizing the inherent limitations of current COD methodologies, which predominantly rely on supervised learning models demanding extensive and accurately annotated datasets, resulting in weak generalization, our research proposes a zero-shot MMCPF that circumvents these challenges. Although MLLMs hold significant potential for broad applications, their effectiveness in COD is hindered and they would make misinterpretations of camouflaged objects. To address this challenge, we further propose a strategic enhancement called the Chain of Visual Perception (CoVP), which significantly improves the perceptual capabilities of MLLMs in camouflaged scenes by leveraging both linguistic and visual cues more effectively. We validate the effectiveness of MMCPF on five widely used COD datasets, containing CAMO, COD10K, NC4K, MoCA-Mask and OV-Camo. Experiments show that MMCPF can outperform all existing state-of-the-art zero-shot COD methods, and achieve competitive performance compared to weakly-supervised and fully-supervised methods, which demonstrates the potential of MMCPF. The Github link of this paper is https://github.com/luckybird1994/MMCPF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d192bc1b-25b4-41b1-b962-043380502e1aCited by top-tier papers10
- Leveraging Hallucinations to Reduce Manual Prompt Dependency in Promptable SegmentationJian Hu, Jiayi Lin, Junchi Yan, Shaogang GongNeurIPS 2024 · 39 citations
- SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement LearningJiaqi Huang, Zunnan Xu, Jun Zhou, Ting Liu et al.NeurIPS 2025 · 33 citations
- Discover, Segment, and Select: A Progressive Mechanism for Zero-shot Camouflaged Object SegmentationYilong Yang, Jianxin Tian, Shengchuan Zhang, Liujuan CaoCVPR 2026 · 3 citations
- Beyond Single Images: Retrieval Self-Augmented Unsupervised Camouflaged Object DetectionJi Du, Xin Wang, Fangwei Hao, Mingyang Yu et al.ICCV 2025 · 2 citations
- Enhancing Prompt Generation with Adaptive Refinement for Camouflaged Object DetectionXuehan Chen, Guangyu Ren, Tianhong Dai, Tania Stathaki et al.ICCV 2025 · 1 citation
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
Related papers
- From Language to Instance: Generative Visual Prompting for Zero-shot Camouflaged Object DetectionZihou Zhang, Hao Li, Zhengwei Yang, Zechao Hu et al.ACM MM 2025 · 1 citation
- Beyond Weak Supervision: MLLMs-Guided Graded Knowledge Distillation for Unsupervised Camouflaged Object DetectionHuafeng Chen, Chenguang Zhu, Yueming Lyu, Caifeng ShanCVPR 2026
- VCoder: Versatile Vision Encoders for Multimodal Large Language ModelsJitesh Jain, Jianwei Yang, Humphrey ShiCVPR 2024
- CGCOD: Class-Guided Camouflaged Object DetectionChenxi Zhang, Qing Zhang, Jiayun Wu, Youwei PangACM MM 2025 · 11 citations
- Language-Driven Cross-Modal Classifier for Zero-Shot Multi-Label Image RecognitionYicheng Liu, Jie Wen, Chengliang Liu, Xiaozhao Fang et al.ICML 2024 · 7 citations
