Multi-Modal Segment Anything Model for Camouflaged Scene Segmentation
Guangyu Ren, Hengyan Liu, Michalis Lazarou, Tania Stathaki
摘要
Camouflaged scenes, where objects blend seamlessly into their environments, pose significant challenges to both human observers and computer vision systems. To address this, we propose a novel framework that leverages off-the-shelf foundation models to generate multi-modal prompts for the Segment Anything Model (SAM), thus eliminating the need for manual prompts and significantly improving overall performance on this downstream task. At first, we generate an image caption using the BLIP model and obtain its text embedding through the use of a text encoder. We then generate a visual embedding through the vision encoder of the BLIP model and use both as inputs to SAM to provide additional semantic information about the image. Finally, we propose a couple of architectural novelties, a) we effectively integrate the multi-modal information in SAM through a multi-level adapter and b) we replace the dense embedding of SAM with the image embedding of its image encoder. Our method achieves new state-of-the-art performance in 11 out of 12 metrics in three benchmark datasets for camouflaged detection. Additionally, our method can be successfully adapted to other tasks such as medical image segmentation performing on par or even outperforming the state-of-the-art methods. Our code is available in https://github.com/ic-qialanqian/Vision-Language-SAM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- ACO-MoE-LoRA: Evolving-while-Training for Adapting Segment Anything Model 2 to Specialized DomainsKaiyi Luo, Bangjun Wang, Li Zhang, Fanzhang Li 等ICML 2026
- Training-Free Open-Vocabulary Camouflaged Object Segmentation via Fine-Grained Object Binding and Adaptive Hybrid PromptPeng Ren, Cheng Jiang, Chuande Yang, Fuming Sun 等CVPR 2026
它引用的顶会 Paper26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Segment Everything Everywhere All at OnceXueyan Zou, Jianwei Yang, Hao Zhang, Feng Li 等NeurIPS 2023 · 被引用 889 次
相关 Paper
- Enhancing Prompt Generation with Adaptive Refinement for Camouflaged Object DetectionXuehan Chen, Guangyu Ren, Tianhong Dai, Tania Stathaki 等ICCV 2025 · 被引用 1 次
- HyperCOD: The First Challenging Benchmark and Baseline for Hyperspectral Camouflaged Object DetectionShuyan Bai, Tingfa Xu, Peifu Liu, Yuhao Qiu 等AAAI 2026
- Endow SAM with Keen Eyes: Temporal-Spatial Prompt Learning for Video Camouflaged Object DetectionWenjun Hui, Zhenfeng Zhu, Shuai Zheng, Yao ZhaoCVPR 2024
- Relax Image-Specific Prompt Requirement in SAM: A Single Generic Prompt for Segmenting Camouflaged ObjectsJian Hu, Jiayi Lin, Shaogang Gong, Weitong CaiAAAI 2024 · 被引用 64 次
- ST-SAM: Multimodal Scene Text Segmentation with Dense Visual and Sparse Textual Prompts via SAMJin Wei, Yaqiang Wu, Jiayi Yan, Zeng Li 等AAAI 2026
