DCEL: Deep Cross-modal Evidential Learning for Text-Based Person Retrieval
Shenshen Li, Xing Xu, Yang Yang, Fumin Shen, Yijun Mo, Yujie Li, Heng Tao Shen
摘要
Text-based person retrieval aims at searching for a pedestrian image from multiple candidates with textual descriptions. It is challenging due to uncertain cross-modal alignments caused by the large intra-class variations. To address the challenge, most existing approaches rely on various attention mechanisms and auxiliary information, yet still struggle with the uncertain cross-modal alignments arising from significant intra-class variation, leading to coarse retrieval results. To this end, we propose a novel framework termed Deep Cross-modal Evidential Learning (DCEL), which deploys evidential deep learning to consider the cross-modal alignment uncertainty. Our DCEL model comprises three components: (1) Bidirectional Evidential Learning, which models alignment uncertainty to measure and mitigate the influence of large intra-class variation; (2) Multi-level Semantic Alignment, which leverages a proposed Semantic Filtration module and image-text similarity distribution to facilitate cross-modal alignments; (3) Cross-modal Relation Learning, which reasons about latent correspondences between multi-level tokens of image and text. Finally, we integrate the advantages of the three proposed components to enhance the model to achieve reliable cross-modal alignments. Our DCEL method consistently outperforms more than ten state-of-the-art methods in supervised, weakly supervised, and domain generalization settings on three benchmarks: CUHK-PEDES, ICFG-PEDES, and RSTPReid.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper8
- Noisy-Correspondence Learning for Text-to-Image Person Re-IdentificationYang Qin, Yingke Chen, Dezhong Peng, Xi Peng 等CVPR 2024 · 被引用 83 次
- Adaptive Uncertainty-Based Learning for Text-Based Person RetrievalShenshen Li, Chen He, Xing Xu, Fumin Shen 等AAAI 2024 · 被引用 59 次
- Prototypical Prompting for Text-to-image Person Re-identificationShuanglin Yan, Jun Liu, Neng Dong, Liyan Zhang 等ACM MM 2024 · 被引用 16 次
- Chat-Driven Text Generation and Interaction for Person RetrievalZequn Xie, Chuxin Wang, Yeqiang Wang, Sihang Cai 等EMNLP 2025 · 被引用 12 次
- UGG-ReID: Uncertainty-Guided Graph Model for Multi-Modal Object Re-IdentificationXixi Wan, Aihua Zheng, Bo Jiang, Beibei Wang 等NeurIPS 2025 · 被引用 4 次
相关 Paper
- DSSL: Deep Surroundings-person Separation Learning for Text-based Person RetrievalAichun Zhu, Zijie Wang, Yifeng Li, Xili Wan 等ACM MM 2021 · 被引用 274 次
- Multi-Modal Disordered Representation Learning Network for Description-Based Person SearchFan Yang, Wei Li, Menglong Yang, Binbin Liang 等AAAI 2024 · 被引用 8 次
- Dual Uncertainty-Guided Feature Alignment Learning for Text-Based Person RetrievalYufei Zheng, Jiawei Liu, Bingyu Hu, Zikun Wei 等ACM MM 2025
- Test-Time Adaptation for Text-Based Person SearchKai Niu, Liucun Shi, Ke Han, Qinzi Zhao 等ACM MM 2025
- Pedestrian-specific Bipartite-aware Similarity Learning for Text-based Person RetrievalFei Shen, Xiangbo Shu, Xiaoyu Du, Jinhui TangACM MM 2023 · 被引用 101 次
