Learning Hierarchical Cross-modal Association with Intra-modal Context for Text-Image Person Retrieval
Yifei Deng, Chenglong Li, Futian Wang, Jin Tang
摘要
Existing Text-Image Person Retrieval (TIPR) methods have made substantial progress in modeling cross-modal associations via contrastive learning frameworks, but usually ignore the fine-grained differences in semantic relevance among different samples, which limits retrieval accuracy. To address this problem, we propose a novel Hierarchical Cross-modal Association framework HCA, which leverages the intra-modal fine-grained semantic relations distilled by single-modal pretrained models to constrain hierarchical cross-modal association between image and text modalities, for accurate TIPR. Specifically, to model hierarchical cross-modal semantic relationships, we propose a Hierarchical Relevance Matching (HRM) module. It partitions the matching strength of image-text pairs by jointly considering identity labels and cross-modal similarity, collaborating with unimodal similarity to construct a hierarchical relevance distribution that serves as a soft supervision signal. HRM not only helps the model better capture varying levels of semantic consistency between image-text pairs but also enhances the overall accuracy of cross-modal association learning. To enhance the ability to capture fine-grained cross-modal semantic relationships, we introduce an Image-guided Ambiguous text Token Modeling (IATM) module. It replaces original tokens with semantically ambiguous ones and leverages image guidance to detect and correct these tokens. This process further improves the fine-grained semantic alignment between images and texts. Experimental results demonstrate that HCA achieves new state-of-the-art performance across multiple datasets, thoroughly validating its effectiveness and advancement in cross-modal retrieval tasks.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Dual-Teacher Interactive Knowledge Distillation Network for Text-to-Visible & Infrared Person RetrievalChenglong Li, Zhengyu Chen, Yifei Deng, Aihua ZhengAAAI 2026
- Cross-modal Fuzzy Alignment Network for Text-Aerial Person Retrieval and A Large-scale BenchmarkYifei Deng, Chenglong Li, Yuyang Zhang, Guyue Hu 等CVPR 2026
- Progressive Multi-cue Alignment for Unaligned RGBT TrackingJiandong Jin, Chenglong Li, Hao Feng, Andong Lu 等CVPR 2026
- Correspondence Cognitive Learning for Multi-Modal Object Re-IdentificationChao Su, Shuying Li, Ruitao Pu, Dezhong Peng 等ICML 2026
相关 Paper
- Cross-Modal Implicit Relation Reasoning and Aligning for Text-to-Image Person RetrievalDing Jiang, Mang YeCVPR 2023
- Adaptive Uncertainty-Based Learning for Text-Based Person RetrievalShenshen Li, Chen He, Xing Xu, Fumin Shen 等AAAI 2024 · 被引用 59 次
- Tackling Alignment Ambiguity in Person Retrieval through Conversational Attribute MiningHao Zou, Runqing Zhang, Jin Ding, xue zhou 等CVPR 2026
- Pedestrian-Centric Discriminative and Fine-grained Semantic Mining for Text-based Person RetrievalYuheng Liang, Haipeng Chen, Yu Liu, Yingda Lyu 等WWW 2026
- Learning Semantic Relationship among Instances for Image-Text MatchingZheren Fu, Zhendong Mao, Yan Song, Yongdong ZhangCVPR 2023
