LapsCore: Language-guided Person Search via Color Reasoning
Yushuang Wu, Zizheng Yan, Xiaoguang Han, Guanbin Li, Changqing Zou, Shuguang Cui
Abstract
The key point of language-guided person search is to construct the cross-modal association between visual and textual input. Existing methods focus on designing multimodal attention mechanisms and novel cross-modal loss functions to learn such association implicitly. We propose a representation learning method for language-guided person search based on color reasoning (LapsCore). It can explicitly build a fine-grained cross-modal association bidirectionally. Specifically, a pair of dual sub-tasks, image colorization and text completion, is designed. In the former task, rich text information is learned to colorize gray images, and the latter one requests the model to understand the image and complete color word vacancies in the captions. The two sub-tasks enable models to learn correct alignments between text phrases and image regions, so that rich multimodal representations can be learned. Extensive experiments on multiple datasets demonstrate the effectiveness and superiority of the proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers27
- Learning Granularity-Unified Representations for Text-to-Image Person Re-identificationZhiyin Shao, Xinyu Zhang, Meng Fang, Zhifeng Lin et al.ACM MM 2022 · 197 citations
- CAIBC: Capturing All-round Information Beyond Color for Text-based Person RetrievalZijie Wang, Aichun Zhu, Jingyi Xue, Xili Wan et al.ACM MM 2022 · 126 citations
- An Empirical Study of CLIP for Text-Based Person SearchMin Cao, Yang Bai, Ziyin Zeng, Mang Ye et al.AAAI 2024 · 111 citations
- PLIP: Language-Image Pre-training for Person Representation LearningJialong Zuo, Jiahao Hong, Feng Zhang, Changqian Yu et al.NeurIPS 2024 · 96 citations
- Noisy-Correspondence Learning for Text-to-Image Person Re-IdentificationYang Qin, Yingke Chen, Dezhong Peng, Xi Peng et al.CVPR 2024 · 83 citations
Builds on7
- CAMP: Cross-Modal Adaptive Message Passing for Text-Image RetrievalZihao Wang, Xihui Liu, Hongsheng Li, Lu Sheng et al.ICCV 2019 · 349 citations
- Adversarial Representation Learning for Text-to-Image MatchingNikolaos Sarafianos, Xiang Xu, Ioannis A. KakadiarisICCV 2019 · 228 citations
- Pose-Guided Multi-Granularity Attention Network for Text-Based Person SearchYa Jing, Chenyang Si, Junbo Wang, Wei Wang et al.AAAI 2020 · 182 citations
- Tag2Pix: Line Art Colorization Using Text Tag With SECat and Changing LossHyunsu Kim, Ho Young Jhoo, Eunhyeok Park, Sungjoo YooICCV 2019 · 119 citations
- Saliency-Guided Attention Network for Image-Sentence MatchingZhong Ji, Haoran Wang, Jungong Han, Yanwei PangICCV 2019 · 96 citations
Related papers
- Cross-Modal Implicit Relation Reasoning and Aligning for Text-to-Image Person RetrievalDing Jiang, Mang YeCVPR 2023
- L-CoDe: Language-Based Colorization Using Color-Object Decoupled ConditionsShuchen Weng, Hao Wu, Zheng Chang, Jiajun Tang et al.AAAI 2022 · 57 citations
- Multi-Modal Disordered Representation Learning Network for Description-Based Person SearchFan Yang, Wei Li, Menglong Yang, Binbin Liang et al.AAAI 2024 · 8 citations
- Adaptive Uncertainty-Based Learning for Text-Based Person RetrievalShenshen Li, Chen He, Xing Xu, Fumin Shen et al.AAAI 2024 · 59 citations
- Prototype-guided Cross-modal Completion and Alignment for Incomplete Text-based Person Re-identificationTiantian Gong, Guodong Du, Junsheng Wang, Yongkang Ding et al.ACM MM 2023 · 9 citations
