Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play Approach
Saehyung Lee, Sangwon Yu, Junsung Park, Jihun Yi, Sungroh Yoon
Abstract
In this paper, we primarily address the issue of dialogue-form context query within the interactive text-to-image retrieval task. Our methodology, PlugIR, actively utilizes the general instruction-following capability of LLMs in two ways. First, by reformulating the dialogueform context, we eliminate the necessity of fine-tuning a retrieval model on existing visual dialogue data, thereby enabling the use of any arbitrary black-box model. Second, we construct the LLM questioner to generate nonredundant questions about the attributes of the target image, based on the information of retrieval candidate images in the current context. This approach mitigates the issues of noisiness and redundancy in the generated questions. Beyond our methodology, we propose a novel evaluation metric, Best log Rank Integral (BRI), for a comprehensive assessment of the interactive retrieval system. PlugIR demonstrates superior performance compared to both zero-shot and fine-tuned baselines in various benchmarks. Additionally, the two methodologies comprising PlugIR can be flexibly applied together or separately in various situations. Our codes are available at https://github.com/ Saehyung-Lee/PlugIR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08b35a29-c88e-4955-80a0-36a0cc867126Cited by top-tier papers17
- ImageScope: Unifying Language-Guided Image Retrieval via Large Multimodal Model Collective ReasoningPengfei Luo, Jingbo Zhou, Tong Xu, Yuan Xia et al.WWW 2025 · 14 citations
- Chat-Driven Text Generation and Interaction for Person RetrievalZequn Xie, Chuxin Wang, Yeqiang Wang, Sihang Cai et al.EMNLP 2025 · 12 citations
- Interactive Cross-modal Learning for Text-3D Scene RetrievalYanglin Feng, Yongxiang Li, Yuan Sun, Yang Qin et al.NeurIPS 2025 · 9 citations
- Diffusion Augmented Retrieval: A Training-Free Approach to Interactive Text-to-Image RetrievalZijun Long, Kangheng Liang, Gerardo Aragon-Camarasa, Richard McCreadie et al.SIGIR 2025 · 6 citations
- Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image RetrievalSiting Li, Xiang Gao, Simon S. DuNeurIPS 2025 · 5 citations
Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
Related papers
- Chatting Makes Perfect: Chat-based Image RetrievalMatan Levy, Rami Ben-Ari, Nir Darshan, Dani LischinskiNeurIPS 2023 · 41 citations
- ReCALL: Recalibrating Capability Degradation for MLLM-based Composed Image RetrievalTianyu Yang, ChenWei He, Xiangzhao Hao, Tianyue Wang et al.CVPR 2026 · 3 citations
- Zero-Shot Composed Image Retrieval via Dual-Stream Instruction-Aware DistillationWenliang Zhong, Robert A. Barton, Weizhi An, Feng Jiang et al.ICCV 2025 · 4 citations
- LLaVA-ReID: Selective Multi-image Questioner for Interactive Person Re-IdentificationYiding Lu, Mouxing Yang, Dezhong Peng, Peng Hu et al.ICML 2025
- Imagine and Seek: Improving Composed Image Retrieval with an Imagined ProxyYou Li, Fan Ma, Yi YangCVPR 2025
