Unifying Multi-Modal Uncertainty Modeling and Semantic Alignment for Text-to-Image Person Re-identification
Zhiwei Zhao, Bin Liu, Yan Lu, Qi Chu, Nenghai Yu
摘要
Text-to-Image person re-identification (TI-ReID) aims to retrieve the images of target identity according to the given textual description. The existing methods in TI-ReID focus on aligning the visual and textual modalities through contrastive feature alignment or reconstructive masked language modeling (MLM). However, these methods parameterize the image/text instances as deterministic embeddings and do not explicitly consider the inherent uncertainty in pedestrian images and their textual descriptions, leading to limited image-text relationship expression and semantic alignment. To address the above problem, in this paper, we propose a novel method that unifies multi-modal uncertainty modeling and semantic alignment for TI-ReID. Specifically, we model the image and textual feature vectors of pedestrian as Gaussian distributions, where the multi-granularity uncertainty of the distribution is estimated by incorporating batch-level and identity-level feature variances for each modality. The multi-modal uncertainty modeling acts as a feature augmentation and provides richer image-text semantic relationship. Then we present a bi-directional cross-modal circle loss to more effectively align the probabilistic features between image and text in a self-paced manner. To further promote more comprehensive image-text semantic alignment, we design a task that complements the masked language modeling, focusing on the cross-modality semantic recovery of global masked token after cross-modal interaction. Extensive experiments conducted on three TI-ReID datasets highlight the effectiveness and superiority of our method over state-of-the-arts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- DM-Adapter: Domain-Aware Mixture-of-Adapters for Text-Based Person RetrievalYating Liu, Zimo Liu, Xiangyuan Lan, Wenming Yang 等AAAI 2025 · 被引用 22 次
- Prototypical Prompting for Text-to-image Person Re-identificationShuanglin Yan, Jun Liu, Neng Dong, Liyan Zhang 等ACM MM 2024 · 被引用 16 次
- Causal Inference over Visual-Semantic-Aligned Graph for Image ClassificationLei Meng, Xiangxian Li, Xiaoshuo Yan, Haokai Ma 等AAAI 2025 · 被引用 11 次
- CLIP-driven View-aware Prompt Learning for Unsupervised Vehicle Re-identificationJiyang Xu, Qi Wang, Xin Xiong, Di Gai 等AAAI 2025 · 被引用 8 次
- Exploring the Potential of Large Vision-Language Models for Unsupervised Text-Based Person RetrievalZongyi Li, Jianbo Li, Yuxuan Shi, Jiazhong Chen 等AAAI 2025 · 被引用 5 次
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li 等AAAI 2020 · 被引用 4,134 次
- DSSL: Deep Surroundings-person Separation Learning for Text-based Person RetrievalAichun Zhu, Zijie Wang, Yifeng Li, Xili Wan 等ACM MM 2021 · 被引用 274 次
- Uncertainty Modeling for Out-of-Distribution GeneralizationXiaotong Li, Yongxing Dai, Yixiao Ge, Jun Liu 等ICLR 2022 · 被引用 237 次
相关 Paper
- KPDM: Key Phrase Dynamic Masking for Robust Text-to-Image Person RetrievalShaofeng You, Tianle Miao, Qihang Chen, Xin Li 等AAAI 2026
- DCEL: Deep Cross-modal Evidential Learning for Text-Based Person RetrievalShenshen Li, Xing Xu, Yang Yang, Fumin Shen 等ACM MM 2023 · 被引用 56 次
- Tackling Alignment Ambiguity in Person Retrieval through Conversational Attribute MiningHao Zou, Runqing Zhang, Jin Ding, xue zhou 等CVPR 2026
- Learning Granularity-Unified Representations for Text-to-Image Person Re-identificationZhiyin Shao, Xinyu Zhang, Meng Fang, Zhifeng Lin 等ACM MM 2022 · 被引用 197 次
- Adaptive Uncertainty-Based Learning for Text-Based Person RetrievalShenshen Li, Chen He, Xing Xu, Fumin Shen 等AAAI 2024 · 被引用 59 次
