LDRE: LLM-based Divergent Reasoning and Ensemble for Zero-Shot Composed Image Retrieval
Zhenyu Yang, Dizhan Xue, Shengsheng Qian, Weiming Dong, Changsheng Xu
Abstract
Zero-Shot Composed Image Retrieval (ZS-CIR) has garnered increasing interest in recent years, which aims to retrieve a target image based on a query composed of a reference image and a modification text without training samples. Specifically, the modification text describes the distinction between the two images. To conduct ZS-CIR, the prevailing methods employ pre-trained image-to-text models to transform the query image and text into a single text, which is then projected into the common feature space by CLIP to retrieve the target image. However, these methods neglect that ZS-CIR is a typicalfuzzy retrieval task, where the semantics of the target image are not strictly defined by the query image and text. To overcome this limitation, this paper proposes a training-free LLM-based Divergent Reasoning and Ensemble (LDRE) method for ZS-CIR to capture diverse possible semantics of the composed result. Firstly, we employ a pre-trained captioning model to generate dense captions for the reference image, focusing on different semantic perspectives of the reference image. Then, we prompt Large Language Models (LLMs) to conduct divergent compositional reasoning based on the dense captions and modification text, deriving divergent edited captions that cover the possible semantics of the composed target. Finally, we design a divergent caption ensemble to obtain the ensemble caption feature weighted by semantic relevance scores, which is subsequently utilized to retrieve the target image in the CLIP feature space. Extensive experiments on three public datasets demonstrate that our proposed LDRE achieves the new state-of-the-art performance.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 604ff4ec-52e9-4742-af2d-63aa6fd69748Cited by top-tier papers32
- LiveStar: Live Streaming Assistant for Real-World Online Video UnderstandingZhenyu Yang, Kairui Zhang, Yuhang Hu, Bing Wang et al.NeurIPS 2025 · 26 citations
- ImageScope: Unifying Language-Guided Image Retrieval via Large Multimodal Model Collective ReasoningPengfei Luo, Jingbo Zhou, Tong Xu, Yuan Xia et al.WWW 2025 · 14 citations
- Leveraging Large Vision-Language Model as User Intent-Aware Encoder for Composed Image RetrievalZelong Sun, Dong Jing, Guoxing Yang, Nanyi Fei et al.AAAI 2025 · 13 citations
- CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image RetrievalZelong Sun, Dong Jing, Zhiwu LuICCV 2025 · 5 citations
- WISER: Wider Search, Deeper Thinking, and Adaptive Fusion for Training-Free Zero-Shot Composed Image RetrievalTianyue Wang, Leigang Qu, Tianyu Yang, Xiangzhao Hao et al.CVPR 2026 · 4 citations
Related papers
- Semantic Editing Increment Benefits Zero-Shot Composed Image RetrievalZhenyu Yang, Shengsheng Qian, Dizhan Xue, Jiahong Wu et al.ACM MM 2024 · 15 citations
- Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image RetrievalYuanmin Tang, Jue Zhang, Xiaoting Qin, Jing Yu et al.CVPR 2025
- SDR-CIR: Semantic Debias Retrieval Framework for Training-Free Zero-Shot Composed Image RetrievalYi Sun, Jinyu Xu, Qing Xie, Jiachen Li et al.WWW 2026 · 1 citation
- G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image RetrievalJiyoung Lim, Heejae Yang, Jee-Hyong LeeCVPR 2026 · 1 citation
- STiTch: Semantic Transition and Transportation in Collaboration for Training-Free Zero-Shot Composed Image RetrievalMiaoge Li, Dongsheng Wang, Zening Sun, Jinsen Zhang et al.CVPR 2026 · 2 citations
