Text-prompt Camouflaged Instance Segmentation with Graduated Camouflage Learning
Zhentao He, Changqun Xia, Shengye Qiao, Jia Li
Abstract
Camouflaged instance segmentation (CIS) aims to detect and segment objects blending with their surroundings. While existing CIS methods rely heavily on fully-supervised training with massive precisely annotated data, consuming considerable annotation efforts yet struggling to segment highly camouflaged objects accurately. Despite their visual similarity to the background, camouflaged objects differ semantically. Since text associated with images offers explicit semantic cues to underscore this difference, we propose a novel approach: the first Text-Prompt based weakly-supervised camouflaged instance segmentation method named TPNet, leveraging semantic distinctions for effective segmentation. TPNet operates in two stages: pseudo mask generation and a self-training process. In the first stage, we align text prompts with images using a language-image model to obtain region proposals containing camouflaged instances. A Semantic-Spatial Iterative Fusion module is designed to assimilate spatial information with semantic insights, iteratively refining pseudo mask. In the second stage, Graduated Camouflage Learning, a self-training strategy, sequences training from simple to complex images based on camouflage levels, facilitating an effective learning gradient. Through the collaboration of the dual phases, our method offers a comprehensive experiment on two common benchmark and demonstrates a significant advancement, delivering a novel solution that bridges the gap between weak-supervised and high camouflaged instance segmentation.
• Computing methodologies → Interest point and salient region detections.
- Correspondence should be addressed to Changqun Xia and Jia Li. We will publish the code and results on the below website. Website: https://cvteam.buaa.edu.cn
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef1acd3d-9916-4c21-9980-b12a6b3d19b9Cited by top-tier papers5
- ST-SAM: SAM-Driven Self-Training Framework for Semi-Supervised Camouflaged Object DetectionXihang Hu, Fuming Sun, Jiazhe Liu, Feilong Xu et al.ACM MM 2025 · 5 citations
- Seeing the Unseen: A Semantic Alignment and Context-Aware Prompt Framework for Open-Vocabulary Camouflaged Object SegmentationPeng Ren, Tian Bai, Jing Sun, Fuming SunICCV 2025 · 4 citations
- Beyond Single Images: Retrieval Self-Augmented Unsupervised Camouflaged Object DetectionJi Du, Xin Wang, Fangwei Hao, Mingyang Yu et al.ICCV 2025 · 2 citations
- Holistic Correction with Object Prototype for Video Object SegmentationShengye Qiao, Changqun Xia, Yanjie Liang, Gongjin Lan et al.AAAI 2025 · 1 citation
- Camouflage-aware Image-Text Retrieval via Expert CollaborationYao Jiang, Zhongkuan Mao, Xuan Wu, Keren Fu et al.CVPR 2026
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Pyramidal Feature Shrinking for Salient Object DetectionMingcan Ma, Changqun Xia, Jia LiAAAI 2021 · 180 citations
Related papers
- Camouflaged Instance Segmentation via Explicit De-CamouflagingNaisong Luo, Yuwen Pan, Rui Sun, Tianzhu Zhang et al.CVPR 2023
- Relax Image-Specific Prompt Requirement in SAM: A Single Generic Prompt for Segmenting Camouflaged ObjectsJian Hu, Jiayi Lin, Shaogang Gong, Weitong CaiAAAI 2024 · 64 citations
- A Unified Query-based Paradigm for Camouflaged Instance SegmentationBo Dong, Jialun Pei, Rongrong Gao, Tian-Zhu Xiang et al.ACM MM 2023 · 19 citations
- CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic SegmentationYuqi Lin, Minghao Chen, Wenxiao Wang, Boxi Wu et al.CVPR 2023
- Weakly-Supervised Text Instance SegmentationXinyan Zu, Haiyang Yu, Bin Li, Xiangyang XueACM MM 2023 · 8 citations
