Rethinking Pseudo Word Learning in Zero-Shot Composed Image Retrieval: From an Object-Aware Perspective
Zhe Li, Lei Zhang, Kun Zhang, Weidong Chen, Yongdong Zhang, Zhendong Mao
Abstract
Composed Image Retrieval (CIR) takes a composed query of a reference image and a text describing the user's intention, with the aim to retrieve the target image under both conditions. Conventional CIR approaches heavily rely on massive annotated triplets, which often comes at a considerable cost. Zero-Shot CIR (ZS-CIR) offers a new solution that can perform diverse CIR tasks without training on the triplet datasets. The key to the ZS-CIR task is to make specified changes to specific objects in the reference image based on the text. Previous works utilize a projection module to map the reference image into single or multiple pseudo words. However, they are either only applicable to single-object scenarios, or naively convert entire image features into multiple pseudo words and fail to focus on the desired target objects specified by the text description. In this work, we rethink how to learn pseudo words based on the objects attended by the text and propose a Multi-Object Aware ZS-CIR framework (MOA). Specifically, a multi-object recognizer first recognizes valid objects in the reference image guided by a set of learnable object queries. Then, we devise an object filtering strategy, which utilizes contextual prompts comprised of noun categories to guide the model in precisely screening out the objects that need to be modified. Finally, the pseudo word learning branch adaptively converts the screened objects into multiple pseudo words for accurate ZS-CIR. Although simple, our MOA consistently outperforms previous state-of-the-art methods across diverse benchmarks and even achieves competitive results with many supervised methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get fc77ba40-3641-4dd0-a06b-3ace483a60a7Cited by top-tier papers3
- WISER: Wider Search, Deeper Thinking, and Adaptive Fusion for Training-Free Zero-Shot Composed Image RetrievalTianyue Wang, Leigang Qu, Tianyu Yang, Xiangzhao Hao et al.CVPR 2026 · 4 citations
- Hierarchy-Aware Pseudo Word Learning with Text Adaptation for Zero-Shot Composed Image RetrievalZhe Li, Lei Zhang, Zheren Fu, Kun Zhang et al.ICCV 2025 · 1 citation
- XR: Cross-Modal Agents for Composed Image RetrievalZhongyu Yang, Wei Pang, Yingfang YuanWWW 2026 · 1 citation
Related papers
- Pic2Word: Mapping Pictures to Words for Zero-shot Composed Image RetrievalKuniaki Saito, Kihyuk Sohn, Xiang Zhang, Chun-Liang Li et al.CVPR 2023
- Fine-grained Textual Inversion Network for Zero-Shot Composed Image RetrievalHaoqiang Lin, Haokun Wen, Xuemeng Song, Meng Liu et al.SIGIR 2024 · 29 citations
- Knowledge-Enhanced Dual-Stream Zero-Shot Composed Image RetrievalYucheng Suo, Fan Ma, Linchao Zhu, Yi YangCVPR 2024 · 20 citations
- Zero-Shot Composed Image Retrieval with Textual InversionAlberto Baldrati, Lorenzo Agnolucci, Marco Bertini, Alberto Del BimboICCV 2023 · 214 citations
- Visual Delta Generator with Large Multi-Modal Models for Semi-Supervised Composed Image RetrievalYoung Kyun Jang, Donghyun Kim, Zihang Meng, Dat Huynh et al.CVPR 2024
