Prototype-guided Cross-modal Completion and Alignment for Incomplete Text-based Person Re-identification
Tiantian Gong, Guodong Du, Junsheng Wang, Yongkang Ding, Liyan Zhang
Abstract
Traditional text-based person re-identification (ReID) techniques heavily rely on fully matched multi-modal data, which is an ideal scenario. However, due to inevitable data missing and corruption during the collection and processing of cross-modal data, the incomplete data issue is usually met in real-world applications. Therefore, we consider a more practical task termed the incomplete text-based ReID task, where person images and text descriptions are not completely matched and contain partially missing modality data. To this end, we propose a novel Prototype-guided Cross-modal Completion and Alignment (PCCA) framework to handle the aforementioned issues for incomplete text-based ReID. Specifically, we cannot directly retrieve person images based on a text query on missing modality data. Therefore, we propose the cross-modal nearest neighbor construction strategy for missing data by computing the cross-modal similarity between existing images and texts, which provides key guidance for the completion of missing modal features. Furthermore, to efficiently complete the missing modal features, we construct the relation graphs with the aforementioned cross-modal nearest neighbor sets of missing modal data and the corresponding prototypes, which can further enhance the generated missing modal features. Additionally, for tighter fine-grained alignment between images and texts, we raise a prototype-aware cross-modal alignment loss that can effectively reduce the modality heterogeneity gap for better fine-grained alignment in common space. Extensive experimental results on several benchmarks with different missing ratios amply demonstrate that our method can consistently outperform state-of-the-art text-image ReID approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6b27e725-9883-4232-8663-be8ee20d7166Cited by top-tier papers1
Ask how each one uses itBuilds on14
- DSSL: Deep Surroundings-person Separation Learning for Text-based Person RetrievalAichun Zhu, Zijie Wang, Yifeng Li, Xili Wan et al.ACM MM 2021 · 274 citations
- Global-Local Temporal Representations for Video Person Re-IdentificationJianing Li, Shiliang Zhang, Jingdong Wang, Wen Gao et al.ICCV 2019 · 241 citations
- Adversarial Representation Learning for Text-to-Image MatchingNikolaos Sarafianos, Xiang Xu, Ioannis A. KakadiarisICCV 2019 · 228 citations
- Learning Granularity-Unified Representations for Text-to-Image Person Re-identificationZhiyin Shao, Xinyu Zhang, Meng Fang, Zhifeng Lin et al.ACM MM 2022 · 197 citations
- Pose-Guided Multi-Granularity Attention Network for Text-Based Person SearchYa Jing, Chenyang Si, Junbo Wang, Wei Wang et al.AAAI 2020 · 182 citations
Related papers
- Cross-Modal Implicit Relation Reasoning and Aligning for Text-to-Image Person RetrievalDing Jiang, Mang YeCVPR 2023
- Text-Based Occluded Person Re-identification via Multi-Granularity Contrastive Consistency LearningXinyi Wu, Wentao Ma, Dan Guo, Tongqing Zhou et al.AAAI 2024 · 29 citations
- Progressive Attribute Embedding for Accurate Cross-modality Person Re-IDAihua Zheng, Peng Pan, Hongchao Li, Chenglong Li et al.ACM MM 2022 · 18 citations
- Weakly Supervised Text-based Person Re-IdentificationShizhen Zhao, Changxin Gao, Yuanjie Shao, Wei-Shi Zheng et al.ICCV 2021 · 39 citations
- Unifying Multi-Modal Uncertainty Modeling and Semantic Alignment for Text-to-Image Person Re-identificationZhiwei Zhao, Bin Liu, Yan Lu, Qi Chu et al.AAAI 2024 · 40 citations
