CamoQuery: Language-Guided Reasoning Camouflaged Object Segmentation
Tianxin Han, Qing Dong, Xingwei Wang, Jie Jia, Gang Wu, Bowen Yang, Fu Zhang
Abstract
Although camouflaged object segmentation has advanced rapidly in recent years, existing methods are still confined to visual mask prediction under fixed task assumptions. They cannot interactively respond to user requests, nor can they proactively understand and reason about the user's intent. Our work tackles this issue by proposing a novel task, Language-Guided Reasoning Camouflaged Object Segmentation (LR-COS). Given a camouflaged image and an implicit query text instruction that requires reasoning, LRCOS aims to output intent-consistent segmentation mask. To establish a benchmark for this task, we build CamoQuery, comprising 12,437 image-mask samples and 25971 implicit query text instructions. To better reflect real-world camouflaged scenarios, we additionally collect MCD, a multi-instance camouflage dataset where multiple camouflaged targets coexist within the same scene, increasing the need for reasoning. Building on CamoQuery, we further propose COSA, a vision-language segmentation assistant that segments the intended camouflaged object from implicit queries and produces a reasoning explanation. Experiments on CamoQuery demonstrate that COSA has strong reasoning segmentation capability in camouflaged scenes and exhibits zero-shot capability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08132fb7-156b-449e-8ecf-014355f30b42Builds on13
- Zoom In and Out: A Mixed-scale Triplet Network for Camouflaged Object DetectionYouwei Pang, Xiaoqi Zhao, Tian-Zhu Xiang, Lihe Zhang et al.CVPR 2022 · 417 citations
- CRIS: CLIP-Driven Referring Image SegmentationZhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao et al.CVPR 2022 · 337 citations
- LAVT: Language-Aware Vision Transformer for Referring Image SegmentationZhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen et al.CVPR 2022 · 319 citations
- Segment, Magnify and Reiterate: Detecting Camouflaged Objects the Hard WayQi Jia, Shuilian Yao, Yu Liu, Xin Fan et al.CVPR 2022 · 230 citations
- GLaMM: Pixel Grounding Large Multimodal ModelHanoona Abdul Rasheed, Muhammad Maaz, Sahal Shaji Mullappilly, Abdelrahman M. Shaker et al.CVPR 2024 · 113 citations
Related papers
- LISA: Reasoning Segmentation via Large Language ModelXin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li et al.CVPR 2024
- Discover, Segment, and Select: A Progressive Mechanism for Zero-shot Camouflaged Object SegmentationYilong Yang, Jianxin Tian, Shengchuan Zhang, Liujuan CaoCVPR 2026 · 3 citations
- Unveiling the Invisible: Reasoning Complex Occlusions Amodally with AURAZhixuan Li, Hyunse Yoon, Sanghoon Lee, Weisi LinICCV 2025
- Training-Free Open-Vocabulary Camouflaged Object Segmentation via Fine-Grained Object Binding and Adaptive Hybrid PromptPeng Ren, Cheng Jiang, Chuande Yang, Fuming Sun et al.CVPR 2026
- Seeing the Unseen: A Semantic Alignment and Context-Aware Prompt Framework for Open-Vocabulary Camouflaged Object SegmentationPeng Ren, Tian Bai, Jing Sun, Fuming SunICCV 2025 · 4 citations
