Chat-based Person Retrieval via Dialogue-Refined Cross-Modal Alignment
Yang Bai, Yucheng Ji, Min Cao, Jinqiao Wang, Mang Ye
摘要
Traditional text-based person retrieval (TPR) relies on a single-shot text as query to retrieve the target person, assuming that the query completely captures the user's search intent. However, in real-world scenarios, it can be challenging to ensure the information completeness of such single-shot text. To address this limitation, we propose chat-based person retrieval (ChatPR), a new paradigm that takes an interactive dialogue as query to perform the person retrieval, engaging the user in conversational context to progressively refine the query for accurate person retrieval. The primary challenge in ChatPR is the lack of available dialogue-image paired data. To overcome this challenge, we establish ChatPedes, the first dataset designed for ChatPR, which is constructed by leveraging large language models to automate the question generation and simulate user responses. Additionally, to bridge the modality gap between dialogues and images, we propose a dialogue-refined cross-modal alignment (DiaNA) framework, which leverages two adaptive attribute refiner modules to bottleneck the conversational and visual information for fine-grained cross-modal alignment. Moreover, we propose a dialoguespecific data augmentation strategy, random round retaining, to further enhance the model's generalization ability across varying dialogue lengths. Extensive experiments demonstrate that DiaNA significantly outperforms existing TPR approaches, highlighting the effectiveness of conversational interactions for person retrieval.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Backdoor Cleaning without External Guidance in MLLM Fine-tuningXuankun Rong, Wenke Huang, Jian Liang, Jinhe Bi 等NeurIPS 2025 · 被引用 39 次
- EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language ModelsYiyang Fang, Wenke Huang, Pei Fu, Yihao Yang 等CVPR 2026 · 被引用 4 次
- Text-based Aerial-Ground Person RetrievalXinyu Zhou, Yu Wu, Jiayao Ma, Wenhao Wang 等AAAI 2026 · 被引用 1 次
- Composite-Attribute Person Re-Identification via Pose-Guided DisentanglementKartik Patwari, Noranart Vesdapunt, Chien-Yi Wang, Dawei Li 等CVPR 2026
- API: Adaptive Prototype Imputation for Incomplete Multimodal Sentiment AnalysisXiaotao Wang, Yiyang Fang, Wenke Huang, Bin Yang 等ICML 2026
它引用的顶会 Paper27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- Interactive Person Retrieval via Multi-Turn Multimodal ConversationYang Bai, Tingfeng Wang, Bin Yang, Min Cao 等ICML 2026
- Tackling Alignment Ambiguity in Person Retrieval through Conversational Attribute MiningHao Zou, Runqing Zhang, Jin Ding, xue zhou 等CVPR 2026
- Dialogue-Driven Interactive Dynamic Learning for Text-to-Image Person RetrievalHongyu Liu, Hongwei Ge, Yuxuan Liu, Yaqing HouACM MM 2025
- GPT-ReID: Learning Fine-grained Representation with GPT for Text-based Person RetrievalXudong Wang, Lei Tan, Pingyang Dai, Liujuan Cao 等ACM MM 2025
- Human-centered Interactive Learning via MLLMs for Text-to-Image Person Re-identificationYang Qin, Chao Chen, Zhihang Fu, Dezhong Peng 等CVPR 2025
