Ask&Confirm: Active Detail Enriching for Cross-Modal Retrieval with Partial Query
Guanyu Cai, Jun Zhang, Xinyang Jiang, Yifei Gong, Lianghua He, Fufu Yu, Pai Peng, Xiaowei Guo, Feiyue Huang, Xing Sun
摘要
Text-based image retrieval has seen considerable progress in recent years. However, the performance of existing methods suffers in real life since the user is likely to provide an incomplete description of an image, which often leads to results filled with false positives that fit the incomplete description. In this work, we introduce the partial-query problem and extensively analyze its influence on text-based image retrieval. Previous interactive methods tackle the problem by passively receiving users’ feedback to supplement the incomplete query iteratively, which is time-consuming and requires heavy user effort. Instead, we propose a novel retrieval framework that conducts the interactive process in an Ask-and-Confirm fashion, where AI actively searches for discriminative details missing in the current query, and users only need to confirm AI’s proposal. Specifically, we propose an object-based interaction to make the interactive retrieval more user-friendly and present a reinforcement-learning-based policy to search for discriminative objects. Furthermore, since fully-supervised training is often infeasible due to the difficulty of obtaining human-machine dialog data, we present a weakly-supervised training strategy that needs no human-annotated dialogs other than a text-image dataset. Experiments show that our framework significantly improves the performance of text-based image retrieval. Code is available at https://github.com/CuthbertCai/Ask-Confirm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Learning to Retrieve Videos by Asking QuestionsAvinash Madasu, Junier Oliva, Gedas BertasiusACM MM 2022 · 被引用 17 次
- Simple Baselines for Interactive Video Retrieval with Questions and AnswersKaiqu Liang, Samuel AlbanieICCV 2023 · 被引用 10 次
- Quantifying and Narrowing the Unknown: Interactive Text-to-Video Retrieval Via Uncertainty MinimizationBingqing Zhang, Zhuo Cao, Heming Du, Yang Li 等ICCV 2025 · 被引用 3 次
- DialogueVPR: Towards Conversational Visual Place RecognitionYukun Song, Changwei Wang, Xingtian Pei, Shibiao Xu 等CVPR 2026 · 被引用 1 次
- Chat-based Person Retrieval via Dialogue-Refined Cross-Modal AlignmentYang Bai, Yucheng Ji, Min Cao, Jinqiao Wang 等CVPR 2025
它引用的顶会 Paper2
相关 Paper
- Conversational Composed Retrieval with Iterative Sequence RefinementHao Wei, Shuhui Wang, Zhe Xue, Shengbo Chen 等ACM MM 2023 · 被引用 4 次
- CAST: Context-Aware Dynamic Latent Space Transformation for Interactive Text-to-Image RetrievalXuanzuo Lin, Min Zhang, Daizong Liu, Zhiwen Zuo 等CVPR 2026
- Chatting Makes Perfect: Chat-based Image RetrievalMatan Levy, Rami Ben-Ari, Nir Darshan, Dani LischinskiNeurIPS 2023 · 被引用 41 次
- Interactive Cross-modal Learning for Text-3D Scene RetrievalYanglin Feng, Yongxiang Li, Yuan Sun, Yang Qin 等NeurIPS 2025 · 被引用 9 次
- Generative Thinking, Corrective Action: User-Friendly Composed Image Retrieval via Automatic Multi-Agent CollaborationZhangtao Cheng, Yuhao Ma, Jian Lang, Kunpeng Zhang 等KDD 2025 · 被引用 2 次
