Hierarchy-Aware Pseudo Word Learning with Text Adaptation for Zero-Shot Composed Image Retrieval
Zhe Li, Lei Zhang, Zheren Fu, Kun Zhang, Zhendong Mao
摘要
Zero-Shot Composed Image Retrieval (ZS-CIR) aims to retrieve the target image based on a reference image and a text describing the user's intention without training on the triplet datasets. The key to this task is to make specified changes to specific objects in the reference image based on the text. Previous works generate single or multiple pseudo words by projecting the reference image to the word embedding space. However, these methods ignore the fact that the editing objects of CIR are naturally hierarchical, and lack the ability of text adaptation, thus failing to adapt to multilevel editing needs. In this paper, we argue that the hierarchical object decomposition is the key to learning pseudo words, and propose a hierarchy-aware dynamic pseudo word learning (HIT) framework to equip with HIerarchy semantic parsing and Text-adaptive filtering. The proposed HIT enjoys several merits. First, HIT is empowered to dynamically decompose the image into different granularity of editing objects by a set of learnable group tokens as guidance, thus naturally forming the hierarchical semantic concepts. Second, the text-adaptive filtering strategy is proposed to screen out specific objects from different levels based on the text, so as to learn hierarchical pseudo words that meet diverse editing needs. Extensive experiments on three challenging benchmarks show that HIT outperforms prior state-of-the-art ones by 5%-8% in average recall.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- GroupViT: Semantic Segmentation Emerges from Text SupervisionJiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon 等CVPR 2022 · 被引用 398 次
- Image Retrieval on Real-life Images with Pre-trained Vision-and-Language ModelsZheyuan Liu, Cristian Rodriguez Opazo, Damien Teney, Stephen GouldICCV 2021 · 被引用 344 次
- Zero-Shot Composed Image Retrieval with Textual InversionAlberto Baldrati, Lorenzo Agnolucci, Marco Bertini, Alberto Del BimboICCV 2023 · 被引用 214 次
相关 Paper
- Rethinking Pseudo Word Learning in Zero-Shot Composed Image Retrieval: From an Object-Aware PerspectiveZhe Li, Lei Zhang, Kun Zhang, Weidong Chen 等SIGIR 2025 · 被引用 5 次
- Knowledge-Enhanced Dual-Stream Zero-Shot Composed Image RetrievalYucheng Suo, Fan Ma, Linchao Zhu, Yi YangCVPR 2024 · 被引用 20 次
- Fine-grained Textual Inversion Network for Zero-Shot Composed Image RetrievalHaoqiang Lin, Haokun Wen, Xuemeng Song, Meng Liu 等SIGIR 2024 · 被引用 29 次
- Pic2Word: Mapping Pictures to Words for Zero-shot Composed Image RetrievalKuniaki Saito, Kihyuk Sohn, Xiang Zhang, Chun-Liang Li 等CVPR 2023
- Context-I2W: Mapping Images to Context-Dependent Words for Accurate Zero-Shot Composed Image RetrievalYuanmin Tang, Jing Yu, Keke Gai, Jiamin Zhuang 等AAAI 2024 · 被引用 71 次
