Human-centered Interactive Learning via MLLMs for Text-to-Image Person Re-identification
Yang Qin, Chao Chen, Zhihang Fu, Dezhong Peng, Xi Peng, Peng Hu
Abstract
Despite remarkable advancements in text-to-image person re-identification (TIReID) facilitated by the breakthrough of cross-modal embedding models, existing methods often struggle to distinguish challenging candidate images due to intrinsic limitations, such as network architecture and data quality. To address these issues, we propose an Interactive Cross-modal Learning framework (ICL), which leverages human-centered interaction to enhance the discriminability of text queries through external multimodal knowledge. To achieve this, we propose a plug-andplay Test-time Humane-centered Interaction (THI) module, which performs visual question answering focused on human characteristics, facilitating multi-round interactions with a multimodal large language model (MLLM) to align query intent with latent target images. Specifically, THI refines user queries based on the MLLM responses to reduce the gap to the best-matching images, thereby boosting ranking accuracy. Additionally, to address the limitation of low-quality training texts, we introduce a novel Reorganization Data Augmentation (RDA) strategy based on information enrichment and diversity enhancement to enhance query discriminability by enriching, decomposing, and reorganizing person descriptions. Extensive experiments on four TIReID benchmarks, i.e., CUHK-PEDES, ICFG-PEDES, RSTPReid, and UFine6926, demonstrate that our method achieves remarkable performance with substantial improvement. Code is available at https://github.com/QinYang79/ICL .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 001c7765-d081-4bb6-994e-61b88ed37e2cCited by top-tier papers7
- Interactive Cross-modal Learning for Text-3D Scene RetrievalYanglin Feng, Yongxiang Li, Yuan Sun, Yang Qin et al.NeurIPS 2025 · 9 citations
- Tackling Alignment Ambiguity in Person Retrieval through Conversational Attribute MiningHao Zou, Runqing Zhang, Jin Ding, xue zhou et al.CVPR 2026
- Composite-Attribute Person Re-Identification via Pose-Guided DisentanglementKartik Patwari, Noranart Vesdapunt, Chien-Yi Wang, Dawei Li et al.CVPR 2026
- MAGIC: Multi-Granularity Language-Informed Image ClusteringXiaohan Zhang, Chao Zhang, Chunlin Chen, Huaxiong LiICML 2026
- Generate, Analyze, and Refine: Training-Free Sound Source Localization via MLLM Meta-ReasoningSubin Park, Jung Uk KimCVPR 2026
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- DSSL: Deep Surroundings-person Separation Learning for Text-based Person RetrievalAichun Zhu, Zijie Wang, Yifeng Li, Xili Wan et al.ACM MM 2021 · 274 citations
- Learning Granularity-Unified Representations for Text-to-Image Person Re-identificationZhiyin Shao, Xinyu Zhang, Meng Fang, Zhifeng Lin et al.ACM MM 2022 · 197 citations
Related papers
- Chat-based Person Retrieval via Dialogue-Refined Cross-Modal AlignmentYang Bai, Yucheng Ji, Min Cao, Jinqiao Wang et al.CVPR 2025
- Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReIDWentao Tan, Changxing Ding, Jiayu Jiang, Fei Wang et al.CVPR 2024 · 34 citations
- Empowering Visible-Infrared Person Re-Identification with Large Foundation ModelsZhangyi Hu, Bin Yang, Mang YeNeurIPS 2024 · 45 citations
- Noisy-Correspondence Learning for Text-to-Image Person Re-IdentificationYang Qin, Yingke Chen, Dezhong Peng, Xi Peng et al.CVPR 2024 · 83 citations
- Chain-of-Thought Guided Multi-Modal Object Re-IdentificationYa Gao, Shihao Li, Zhaojun Liu, Aihua Zheng et al.CVPR 2026
