Dual Uncertainty-Guided Feature Alignment Learning for Text-Based Person Retrieval
Yufei Zheng, Jiawei Liu, Bingyu Hu, Zikun Wei, Yong Wu, Zheng-Jun Zha
摘要
Text-based person retrieval (TBPR) aims to retrieve pedestrian images based on textual descriptions, facing challenges due to the inherent heterogeneity and uncertainty between visual and textual modalities. Most existing methods focus on addressing heterogeneity while neglecting the issue of uncertainty. To tackle the uncertainty arising from the diverse textual expressions, including both structural and semantic content variations, we propose a novel Dual Uncertainty-Guided Feature Alignment Learning (DUAL) approach, utilizing instance-level and identity-level uncertainty estimations to mitigate these impacts. Specifically, for the uncertainty caused by textual structure variations, DUAL first introduces an uncertainty Gaussian modeling module that represents image and text features as Gaussian distributions in a learnable manner, and estimates instance-level uncertainty coefficients to quantify structural differences within the text. Subsequently, DUAL leverages ShareGPT4V to standardize the text structure, dynamically aligning the original text features with structure-invariant generated text features through adaptive knowledge distillation guided by the instance-level uncertainty coefficients, effectively reducing structural diversity's impact while minimizing noise. Moreover, for the uncertainty caused by the diversity of textual semantic content, DUAL designs an alignment loss that utilizes identity-level uncertainty coefficients, estimated via a Gaussian Mixture Model based on the distances between image and text features of the same identity, effectively mitigating the impact of semantic content diversity. Experimental results demonstrate that DUAL outperforms existing methods on TBPR benchmarks, highlighting its superiority in multimodal person retrieval.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Adaptive Uncertainty-Based Learning for Text-Based Person RetrievalShenshen Li, Chen He, Xing Xu, Fumin Shen 等AAAI 2024 · 被引用 59 次
- DCEL: Deep Cross-modal Evidential Learning for Text-Based Person RetrievalShenshen Li, Xing Xu, Yang Yang, Fumin Shen 等ACM MM 2023 · 被引用 56 次
- Exploring the Potential of Large Vision-Language Models for Unsupervised Text-Based Person RetrievalZongyi Li, Jianbo Li, Yuxuan Shi, Jiazhong Chen 等AAAI 2025 · 被引用 5 次
- Pedestrian-Centric Discriminative and Fine-grained Semantic Mining for Text-based Person RetrievalYuheng Liang, Haipeng Chen, Yu Liu, Yingda Lyu 等WWW 2026
- Unifying Multi-Modal Uncertainty Modeling and Semantic Alignment for Text-to-Image Person Re-identificationZhiwei Zhao, Bin Liu, Yan Lu, Qi Chu 等AAAI 2024 · 被引用 40 次
