Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval
Haoqiang Lin, Haokun Wen, Xuemeng Song, Meng Liu, Yupeng Hu, Liqiang Nie
摘要
Composed Image Retrieval (CIR) allows users to search target images with a multimodal query, comprising a reference image and a modification text that describes the user's modification demand over the reference image. Nevertheless, due to the expensive labor cost of training data annotation, recent researchers have shifted to the challenging task of zero-shot CIR (ZS-CIR), which targets fulfilling CIR without annotated triplets. The pioneer ZS-CIR studies focus on converting the CIR task into a standard text-to-image retrieval task by pre-training a textual inversion network that can map a given image into a single pseudo-word token. Despite their significant progress, their coarse-grained textual inversion may be insufficient to capture the full content of the image accurately. To overcome this issue, in this work, we propose a novel Fine-grained Textual Inversion Network for ZS-CIR, named FTI4CIR. In particular, FTI4CIR comprises two main components: fine-grained pseudo-word token mapping and tri-wise caption-based semantic regularization. The former maps the image into a subject-oriented pseudo-word token and several attribute-oriented pseudo-word tokens to comprehensively express the image in the textual form, while the latter works on jointly aligning the fine-grained pseudo-word tokens to the real-word token embedding space based on a BLIP-generated image caption template. Extensive experiments conducted on three benchmark datasets demonstrate the superiority of our proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Simple but Effective Raw-Data Level Multimodal Fusion for Composed Image RetrievalHaokun Wen, Xuemeng Song, Xiaolin Chen, Yinwei Wei 等SIGIR 2024 · 被引用 30 次
- ImageScope: Unifying Language-Guided Image Retrieval via Large Multimodal Model Collective ReasoningPengfei Luo, Jingbo Zhou, Tong Xu, Yuan Xia 等WWW 2025 · 被引用 14 次
- Modeling Uncertainty in Composed Image Retrieval via Probabilistic EmbeddingsHaomiao Tang, Jinpeng Wang, Yuang Peng, Guanghao Meng 等ACL 2025 · 被引用 8 次
- HUD: Hierarchical Uncertainty-Aware Disambiguation Network for Composed Video RetrievalZhiwei Chen, Yupeng Hu, Zixu Li, Zhiheng Fu 等ACM MM 2025 · 被引用 5 次
- FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image RetrievalBohan Hou, Haoqiang Lin, Xuemeng Song, Haokun Wen 等SIGIR 2025 · 被引用 2 次
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- SimVLM: Simple Visual Language Model Pretraining with Weak SupervisionZirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai 等ICLR 2022 · 被引用 950 次
- Unifying Vision-and-Language Tasks via Text GenerationJaemin Cho, Jie Lei, Hao Tan, Mohit BansalICML 2021 · 被引用 624 次
相关 Paper
- Rethinking Pseudo Word Learning in Zero-Shot Composed Image Retrieval: From an Object-Aware PerspectiveZhe Li, Lei Zhang, Kun Zhang, Weidong Chen 等SIGIR 2025 · 被引用 5 次
- Pic2Word: Mapping Pictures to Words for Zero-shot Composed Image RetrievalKuniaki Saito, Kihyuk Sohn, Xiang Zhang, Chun-Liang Li 等CVPR 2023
- Zero-Shot Composed Image Retrieval with Textual InversionAlberto Baldrati, Lorenzo Agnolucci, Marco Bertini, Alberto Del BimboICCV 2023 · 被引用 214 次
- Hierarchy-Aware Pseudo Word Learning with Text Adaptation for Zero-Shot Composed Image RetrievalZhe Li, Lei Zhang, Zheren Fu, Kun Zhang 等ICCV 2025 · 被引用 1 次
- Modality and Task Adaptation for Enhanced Zero-shot Composed Image RetrievalHaiwen Li, Delong Liu, Zhaohui Hou, Zeliang Ma 等AAAI 2026 · 被引用 1 次
