Multi-Modal Segment Anything Model for Camouflaged Scene Segmentation
Guangyu Ren, Hengyan Liu, Michalis Lazarou, Tania Stathaki
Abstract
Camouflaged scenes, where objects blend seamlessly into their environments, pose significant challenges to both human observers and computer vision systems. To address this, we propose a novel framework that leverages off-the-shelf foundation models to generate multi-modal prompts for the Segment Anything Model (SAM), thus eliminating the need for manual prompts and significantly improving overall performance on this downstream task. At first, we generate an image caption using the BLIP model and obtain its text embedding through the use of a text encoder. We then generate a visual embedding through the vision encoder of the BLIP model and use both as inputs to SAM to provide additional semantic information about the image. Finally, we propose a couple of architectural novelties, a) we effectively integrate the multi-modal information in SAM through a multi-level adapter and b) we replace the dense embedding of SAM with the image embedding of its image encoder. Our method achieves new state-of-the-art performance in 11 out of 12 metrics in three benchmark datasets for camouflaged detection. Additionally, our method can be successfully adapted to other tasks such as medical image segmentation performing on par or even outperforming the state-of-the-art methods. Our code is available in https://github.com/ic-qialanqian/Vision-Language-SAM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ddb6624f-04cd-4474-87a3-4e1cea73e1a7Cited by top-tier papers2
- ACO-MoE-LoRA: Evolving-while-Training for Adapting Segment Anything Model 2 to Specialized DomainsKaiyi Luo, Bangjun Wang, Li Zhang, Fanzhang Li et al.ICML 2026
- Training-Free Open-Vocabulary Camouflaged Object Segmentation via Fine-Grained Object Binding and Adaptive Hybrid PromptPeng Ren, Cheng Jiang, Chuande Yang, Fuming Sun et al.CVPR 2026
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Segment Everything Everywhere All at OnceXueyan Zou, Jianwei Yang, Hao Zhang, Feng Li et al.NeurIPS 2023 · 889 citations
Related papers
- Enhancing Prompt Generation with Adaptive Refinement for Camouflaged Object DetectionXuehan Chen, Guangyu Ren, Tianhong Dai, Tania Stathaki et al.ICCV 2025 · 1 citation
- HyperCOD: The First Challenging Benchmark and Baseline for Hyperspectral Camouflaged Object DetectionShuyan Bai, Tingfa Xu, Peifu Liu, Yuhao Qiu et al.AAAI 2026
- Endow SAM with Keen Eyes: Temporal-Spatial Prompt Learning for Video Camouflaged Object DetectionWenjun Hui, Zhenfeng Zhu, Shuai Zheng, Yao ZhaoCVPR 2024
- Relax Image-Specific Prompt Requirement in SAM: A Single Generic Prompt for Segmenting Camouflaged ObjectsJian Hu, Jiayi Lin, Shaogang Gong, Weitong CaiAAAI 2024 · 64 citations
- ST-SAM: Multimodal Scene Text Segmentation with Dense Visual and Sparse Textual Prompts via SAMJin Wei, Yaqiang Wu, Jiayi Yan, Zeng Li et al.AAAI 2026
