Camouflage-aware Image-Text Retrieval via Expert Collaboration
Yao Jiang, Zhongkuan Mao, Xuan Wu, Keren Fu, Qijun Zhao
Abstract
Camouflaged scene understanding (CSU) has attracted significant attention due to its broad practical implications. However, in this field, robust image-text cross-modal alignment remains under-explored, hindering deeper understanding of camouflaged scenarios and their related applications. To this end, we focus on the typical image-text retrieval task, and formulate a new task dubbed ``camouflage-aware image-text retrieval'' (CA-ITR). We first construct a dedicated camouflage image-text retrieval dataset (CamoIT), comprising 10.5K samples with multi-granularity textual annotations. Benchmark results conducted on CamoIT reveal the underlying challenges of CA-ITR for existing cutting-edge retrieval techniques, which are mainly caused by objects' camouflage properties as well as those complex image contents. As a solution, we propose a camouflage-expert collaborative network (CECNet), which features a dual-branch visual encoder: one branch captures holistic image representations, while the other incorporates a dedicated model to inject representations of camouflaged objects. A novel confidence-conditioned graph attention (C2GA) mechanism is incorporated to exploit the complementarity across branches. Comparative experiments show that CECNet achieves 29% overall CA-ITR accuracy boost, surpassing seven representative retrieval models. The dataset and code will be available at https://github.com/jiangyao-scu/CA-ITR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dae6fdbd-4c98-4a39-aebc-38c1e006ae46Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- Zoom In and Out: A Mixed-scale Triplet Network for Camouflaged Object DetectionYouwei Pang, Xiaoqi Zhao, Tian-Zhu Xiang, Lihe Zhang et al.CVPR 2022 · 417 citations
- Dual-level Collaborative Transformer for Image CaptioningYunpeng Luo, Jiayi Ji, Xiaoshuai Sun, Liujuan Cao et al.AAAI 2021 · 349 citations
- CAMP: Cross-Modal Adaptive Message Passing for Text-Image RetrievalZihao Wang, Xihui Liu, Hongsheng Li, Lu Sheng et al.ICCV 2019 · 349 citations
Related papers
- CGCOD: Class-Guided Camouflaged Object DetectionChenxi Zhang, Qing Zhang, Jiayun Wu, Youwei PangACM MM 2025 · 11 citations
- Joint Attribute Manipulation and Modality Alignment Learning for Composing Text and Image to Image RetrievalFeifei Zhang, Mingliang Xu, Qirong Mao, Changsheng XuACM MM 2020 · 39 citations
- MiNet: Weakly-Supervised Camouflaged Object Detection through Mutual Interaction between Region and Edge CuesYuzhen Niu, Lifen Yang, Rui Xu, Yuezhou Li et al.ACM MM 2024 · 14 citations
- Context-Aware Attention Network for Image-Text RetrievalQi Zhang, Zhen Lei, Zhaoxiang Zhang, Stan Z. LiCVPR 2020
- High-Resolution Iterative Feedback Network for Camouflaged Object DetectionXiaobin Hu, Shuo Wang, Xuebin Qin, Hang Dai et al.AAAI 2023 · 236 citations
