Simple but Effective Raw-Data Level Multimodal Fusion for Composed Image Retrieval
Haokun Wen, Xuemeng Song, Xiaolin Chen, Yinwei Wei, Liqiang Nie, Tat-Seng Chua
Abstract
Composed image retrieval (CIR) aims to retrieve the target image based on a multimodal query, i.e., a reference image paired with corresponding modification text. Recent CIR studies leverage visionlanguage pre-trained (VLP) methods as the feature extraction backbone, and perform nonlinear feature-level multimodal query fusion to retrieve the target image. Despite the promising performance, we argue that their nonlinear feature-level multimodal fusion may lead to the fused feature deviating from the original embedding space, potentially hurting the retrieval performance. To address this issue, in this work, we propose shifting the multimodal fusion from the feature level to the raw-data level to fully exploit the VLP model's multimodal encoding and cross-modal alignment abilities. In particular, we introduce a Dual Query Unification-based Composed Image Retrieval framework (DQU-CIR), whose backbone simply involves a VLP model's image encoder and a text encoder. Specifically, DQU-CIR first employs two training-free query unification components: text-oriented query unification and vision-oriented query unification, to derive a unified textual and visual query based on the raw data of the multimodal query, respectively. The unified textual query is derived by concatenating the modification text with the extracted reference image's textual description, while the unified visual query is created by writing the key modification words onto the reference image. Ultimately, to address diverse search intentions, DQU-CIR linearly combines the features of the two unified queries encoded by the VLP model to retrieve the target image. Extensive experiments on four real-world datasets validate the effectiveness of our proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1eeafe7c-cda2-4a12-ae02-f3b4c01dd878Cited by top-tier papers12
- OFFSET: Segmentation-based Focus Shift Revision for Composed Image RetrievalZhiwei Chen, Yupeng Hu, Zixu Li, Zhiheng Fu et al.ACM MM 2025 · 10 citations
- Mitigating Modality Bias in Multi-modal Entity Alignment from a Causal PerspectiveTaoyu Su, Jiawei Sheng, Duohe Ma, Xiaodong Li et al.SIGIR 2025 · 4 citations
- WISER: Wider Search, Deeper Thinking, and Adaptive Fusion for Training-Free Zero-Shot Composed Image RetrievalTianyue Wang, Leigang Qu, Tianyu Yang, Xiangzhao Hao et al.CVPR 2026 · 4 citations
- FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image RetrievalBohan Hou, Haoqiang Lin, Xuemeng Song, Haokun Wen et al.SIGIR 2025 · 2 citations
- Hierarchy-Aware Pseudo Word Learning with Text Adaptation for Zero-Shot Composed Image RetrievalZhe Li, Lei Zhang, Zheren Fu, Kun Zhang et al.ICCV 2025 · 1 citation
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- DiffusionCLIP: Text-Guided Diffusion Models for Robust Image ManipulationGwanghyun Kim, Taesung Kwon, Jong Chul YeCVPR 2022 · 458 citations
Related papers
- Leveraging Large Vision-Language Model as User Intent-Aware Encoder for Composed Image RetrievalZelong Sun, Dong Jing, Guoxing Yang, Nanyi Fei et al.AAAI 2025 · 13 citations
- Instance-Level Composed Image RetrievalBill Psomas, George Retsinas, Nikos Efthymiadis, Panagiotis Paraskevas Filntisis et al.NeurIPS 2025 · 14 citations
- Target-Guided Composed Image RetrievalHaokun Wen, Xian Zhang, Xuemeng Song, Yinwei Wei et al.ACM MM 2023 · 53 citations
- Sentence-level Prompts Benefit Composed Image RetrievalYang Bai, Xinxing Xu, Yong Liu, Salman Khan et al.ICLR 2024 · 75 citations
- CoLLM: A Large Language Model for Composed Image RetrievalChuong Huynh, Jinyu Yang, Ashish Tawari, Mubarak Shah et al.CVPR 2025
