CAST: Context-Aware Dynamic Latent Space Transformation for Interactive Text-to-Image Retrieval
Xuanzuo Lin, Min Zhang, Daizong Liu, Zhiwen Zuo, Xun Yang, Changting Lin, Xun Wang, Jianfeng Dong
2026Year
Abstract
Interactive Text-to-Image Retrieval (I-TIR) aims to refine image retrieval results through natural language dialogues, which allows users to progressively supplement or correct their search intention across multiple rounds, enabling a more precise and user-aligned visual search experience.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0db18d46-abe0-4ad0-b6e8-5e00ae1cf3f3Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- Conversational Composed Retrieval with Iterative Sequence RefinementHao Wei, Shuhui Wang, Zhe Xue, Shengbo Chen et al.ACM MM 2023 · 4 citations
- GenIR: Generative Visual Feedback for Mental Image RetrievalDiji Yang, Minghao Liu, Chung-Hsiang Lo, Yi Zhang et al.NeurIPS 2025 · 4 citations
- Dialogue-Driven Interactive Dynamic Learning for Text-to-Image Person RetrievalHongyu Liu, Hongwei Ge, Yuxuan Liu, Yaqing HouACM MM 2025
- Diffusion Augmented Retrieval: A Training-Free Approach to Interactive Text-to-Image RetrievalZijun Long, Kangheng Liang, Gerardo Aragon-Camarasa, Richard McCreadie et al.SIGIR 2025 · 6 citations
- Chatting Makes Perfect: Chat-based Image RetrievalMatan Levy, Rami Ben-Ari, Nir Darshan, Dani LischinskiNeurIPS 2023 · 41 citations
