Cross-Modal Implicit Relation Reasoning and Aligning for Text-to-Image Person Retrieval
Ding Jiang, Mang Ye
摘要
Text-to-image person retrieval aims to identify the target person based on a given textual description query. The primary challenge is to learn the mapping of visual and textual modalities into a common latent space. Prior works have attempted to address this challenge by leveraging separately pre-trained unimodal models to extract visual and textual features. However, these approaches lack the necessary underlying alignment capabilities required to match multimodal data effectively. Besides, these works use prior information to explore explicit part alignments, which may lead to the distortion of intra-modality information. To alleviate these issues, we present IRRA: a cross-modal Implicit Relation Reasoning and Aligning framework that learns relations between local visual-textual tokens and enhances global image-text matching without requiring additional prior supervision. Specifically, we first design an Implicit Relation Reasoning module in a masked language modeling paradigm. This achieves cross-modal interaction by integrating the visual cues into the textual tokens with a cross-modal multimodal interaction encoder. Secondly, to globally align the visual and textual embeddings, Similarity Distribution Matching is proposed to minimize the KL divergence between image-text similarity distributions and the normalized label matching distributions. The proposed method achieves new state-of-the-art results on all three public datasets, with a notable margin of about 3%-9% for Rank-1 accuracy compared to prior methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper79
- An Empirical Study of CLIP for Text-Based Person SearchMin Cao, Yang Bai, Ziyin Zeng, Mang Ye 等AAAI 2024 · 被引用 111 次
- PLIP: Language-Image Pre-training for Person Representation LearningJialong Zuo, Jiahao Hong, Feng Zhang, Changqian Yu 等NeurIPS 2024 · 被引用 96 次
- ChatTS: Aligning Time Series with LLMs via Synthetic Data for Enhanced Understanding and ReasoningZhe Xie, Zeyan Li, Xiao He, Longlong Xu 等VLDB 2025 · 被引用 87 次
- Noisy-Correspondence Learning for Text-to-Image Person Re-IdentificationYang Qin, Yingke Chen, Dezhong Peng, Xi Peng 等CVPR 2024 · 被引用 83 次
- Adaptive Uncertainty-Based Learning for Text-Based Person RetrievalShenshen Li, Chen He, Xing Xu, Fumin Shen 等AAAI 2024 · 被引用 59 次
它引用的顶会 Paper19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
相关 Paper
- Learning Hierarchical Cross-modal Association with Intra-modal Context for Text-Image Person RetrievalYifei Deng, Chenglong Li, Futian Wang, Jin TangACM MM 2025 · 被引用 2 次
- Tackling Alignment Ambiguity in Person Retrieval through Conversational Attribute MiningHao Zou, Runqing Zhang, Jin Ding, xue zhou 等CVPR 2026
- DCEL: Deep Cross-modal Evidential Learning for Text-Based Person RetrievalShenshen Li, Xing Xu, Yang Yang, Fumin Shen 等ACM MM 2023 · 被引用 56 次
- Test-Time Adaptation for Text-Based Person SearchKai Niu, Liucun Shi, Ke Han, Qinzi Zhao 等ACM MM 2025
- Cross-modal Joint Prediction and Alignment for Composed Query Image RetrievalYuchen Yang, Min Wang, Wengang Zhou, Houqiang LiACM MM 2021 · 被引用 29 次
