Beyond Weak Supervision: MLLMs-Guided Graded Knowledge Distillation for Unsupervised Camouflaged Object Detection
Huafeng Chen, Chenguang Zhu, Yueming Lyu, Caifeng Shan
Abstract
Most Camouflaged Object Detection (COD) methods rely on costly pixel-level annotations. Recent studies have adopted unsupervised COD (UCOD) to eliminate labeling costs, but still suffer from two issues:1) insufficient supervision, leading to reliance on self-supervised backbone DINO and reduced model flexibility; and 2) ineffective use of pseudo-labels, which widens the performance gap with supervised methods and limits real-world applicability. In this paper, we propose a novel teacher-student framework for UCOD to address these two issues. To tackle the lack of supervision, we build a powerful teacher model by integrating Multimodal Large Language Models (MLLMs) and the Segment Anything Model (SAM) to generate high-quality pseudo-labels. However, the teacher model faces two challenges: 1) suboptimal performance of MLLMs in COD, and 2) cascading errors.To address these challenges, we first propose a Camouflaged-Aware Chain-of-Thought (CA-CoT) for MLLMs. CA-CoT guides MLLMs through step-by-step reasoning to simulate human perceptual processes, thereby enhancing their performance in COD.Subsequently, we design a Graded Mask Evaluator (GME) to mitigate cascading errors, which evaluates and grades the quality of masks generated by SAM, and then filters out the low-quality masks to provide more reliable supervision.To better leverage pseudo-labels, we propose Graded Knowledge Distillation (GKD), which adaptively enhances distillation at both image and pixel levels based on pseudo-label quality.Extensive experiments show that our method outperforms existing UCOD approaches by a large margin and achieves performance comparable to weakly supervised methods. Notably, our method also achieves good performance under zero-shot settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 201411f2-a340-4d23-a309-d98fe696b154Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Zoom In and Out: A Mixed-scale Triplet Network for Camouflaged Object DetectionYouwei Pang, Xiaoqi Zhao, Tian-Zhu Xiang, Lihe Zhang et al.CVPR 2022 · 417 citations
Related papers
- UCOD-DPL: Unsupervised Camouflaged Object Detection via Dynamic Pseudo-label LearningWeiqi Yan, Lvhai Chen, Huaijia Kou, Shengchuan Zhang et al.CVPR 2025
- ST-SAM: SAM-Driven Self-Training Framework for Semi-Supervised Camouflaged Object DetectionXihang Hu, Fuming Sun, Jiazhe Liu, Feilong Xu et al.ACM MM 2025 · 5 citations
- Exploring Deeper! Segment Anything Model with Depth Perception for Camouflaged Object DetectionZhenni Yu, Xiaoqin Zhang, Li Zhao, Yi Bin et al.ACM MM 2024 · 41 citations
- Relax Image-Specific Prompt Requirement in SAM: A Single Generic Prompt for Segmenting Camouflaged ObjectsJian Hu, Jiayi Lin, Shaogang Gong, Weitong CaiAAAI 2024 · 64 citations
- Improving SAM for Camouflaged Object Detection via Dual Stream AdaptersJiaming Liu, Linghe Kong, Guihai ChenICCV 2025 · 5 citations
