Chatting Makes Perfect: Chat-based Image Retrieval
Matan Levy, Rami Ben-Ari, Nir Darshan, Dani Lischinski
摘要
Chats emerge as an effective user-friendly approach for information retrieval, and are successfully employed in many domains, such as customer service, healthcare, and finance. However, existing image retrieval approaches typically address the case of a single query-to-image round, and the use of chats for image retrieval has been mostly overlooked. In this work, we introduce ChatIR: a chat-based image retrieval system that engages in a conversation with the user to elicit information, in addition to an initial query, in order to clarify the user's search intent. Motivated by the capabilities of today's foundation models, we leverage Large Language Models to generate follow-up questions to an initial image description. These questions form a dialog with the user in order to retrieve the desired image from a large corpus. In this study, we explore the capabilities of such a system tested on a large dataset and reveal that engaging in a dialog yields significant gains in image retrieval. We start by building an evaluation pipeline from an existing manually generated dataset and explore different modules and training strategies for ChatIR. Our comparison includes strong baselines derived from related applications trained with Reinforcement Learning. Our system is capable of retrieving the target image from a pool of 50K images with over 78% success rate after 5 dialogue rounds, compared to 75% when questions are asked by humans, and 64% for a single shot text-to-image retrieval. Extensive evaluations reveal the strong capabilities and examine the limitations of CharIR under different settings. Project repository is available at https://github.com/levymsn/ChatIR .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Vision-by-Language for Training-Free Compositional Image RetrievalShyamgopal Karthik, Karsten Roth, Massimiliano Mancini, Zeynep AkataICLR 2024 · 被引用 120 次
- Memory Reviver: Supporting Photo-Collection Reminiscence for People with Visual Impairment via a Proactive ChatbotShuchang Xu, Chang Chen, Zichen Liu, Xiaofu Jin 等UIST 2024 · 被引用 21 次
- MUST: An Effective and Scalable Framework for Multimodal Search of Target ModalityMengzhao Wang, Xiangyu Ke, Xiaoliang Xu, Lu Chen 等ICDE 2024 · 被引用 16 次
- ImageScope: Unifying Language-Guided Image Retrieval via Large Multimodal Model Collective ReasoningPengfei Luo, Jingbo Zhou, Tong Xu, Yuan Xia 等WWW 2025 · 被引用 14 次
- Interactive Cross-modal Learning for Text-3D Scene RetrievalYanglin Feng, Yongxiang Li, Yuan Sun, Yang Qin 等NeurIPS 2025 · 被引用 9 次
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play ApproachSaehyung Lee, Sangwon Yu, Junsung Park, Jihun Yi 等ACL 2024 · 被引用 7 次
- Diffusion Augmented Retrieval: A Training-Free Approach to Interactive Text-to-Image RetrievalZijun Long, Kangheng Liang, Gerardo Aragon-Camarasa, Richard McCreadie 等SIGIR 2025 · 被引用 6 次
- ChatRetriever: Adapting Large Language Models for Generalized and Robust Conversational Dense RetrievalKelong Mao, Chenlong Deng, Haonan Chen, Fengran Mo 等EMNLP 2024 · 被引用 7 次
- CAST: Context-Aware Dynamic Latent Space Transformation for Interactive Text-to-Image RetrievalXuanzuo Lin, Min Zhang, Daizong Liu, Zhiwen Zuo 等CVPR 2026
- GenIR: Generative Visual Feedback for Mental Image RetrievalDiji Yang, Minghao Liu, Chung-Hsiang Lo, Yi Zhang 等NeurIPS 2025 · 被引用 4 次
