Dialogue-Driven Interactive Dynamic Learning for Text-to-Image Person Retrieval
Hongyu Liu, Hongwei Ge, Yuxuan Liu, Yaqing Hou
Abstract
Text-to-image person retrieval aims to identify target person images using natural language descriptions. Current state-of-the-art methods predominantly rely on single-round retrieval frameworks, where retrieval accuracy heavily depends on the quality of the initial textual descriptions. However, users sometimes struggle to provide detailed and distinctive descriptions in a single attempt, resulting in generic initial queries that lack discriminative details. This fundamental limitation of the single-round retrieval framework frequently leads to the misinterpretation of user intent and suboptimal retrieval performance. To address this limitation, we propose Dialogue-driven Interactive Dynamic Learning (DIDL) for text-to-image person retrieval. Specifically, we first introduce Collaborative Query Refinement (CQR), which progressively refines retrieval conditions through multi-round dialogues. Then, we design Dynamic Context Resampling (DCR) based on a bi-granular mask strategy that enhances the model's adaptation to dialogue-style contexts and effectively balances its attention between initial descriptions and supplementary information. Based on these components, we further propose cross-modal Probabilistic Context Matching Modeling (ProCMM) that establishes effective associations between static visual features and dynamic contextual semantics. Extensive experiments demonstrate that our approach achieves state-of-the-art performance across all three benchmark datasets.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 18dc5d4e-5629-4bf2-862c-f8a0f3a1499fRelated papers
- Chat-based Person Retrieval via Dialogue-Refined Cross-Modal AlignmentYang Bai, Yucheng Ji, Min Cao, Jinqiao Wang et al.CVPR 2025
- Conversational Composed Retrieval with Iterative Sequence RefinementHao Wei, Shuhui Wang, Zhe Xue, Shengbo Chen et al.ACM MM 2023 · 4 citations
- CAST: Context-Aware Dynamic Latent Space Transformation for Interactive Text-to-Image RetrievalXuanzuo Lin, Min Zhang, Daizong Liu, Zhiwen Zuo et al.CVPR 2026
- Tackling Alignment Ambiguity in Person Retrieval through Conversational Attribute MiningHao Zou, Runqing Zhang, Jin Ding, xue zhou et al.CVPR 2026
- Learning Hierarchical Cross-modal Association with Intra-modal Context for Text-Image Person RetrievalYifei Deng, Chenglong Li, Futian Wang, Jin TangACM MM 2025 · 2 citations
