AXM-Net: Implicit Cross-Modal Feature Alignment for Person Re-identification
Ammarah Farooq, Muhammad Awais, Josef Kittler, Syed Safwan Khalid
Abstract
Cross-modal person re-identification (Re-ID) is critical for modern video surveillance systems. The key challenge is to align cross-modality representations induced by the semantic information present for a person and ignore background information. This work presents a novel convolutional neural network (CNN) based architecture designed to learn semantically aligned cross-modal visual and textual representations. The underlying building block, named AXM-Block, is a unified multi-layer network that dynamically exploits the multi-scale knowledge from both modalities and re-calibrates each modality according to shared semantics. To complement the convolutional design, contextual attention is applied in the text branch to manipulate long-term dependencies. Moreover, we propose a unique design to enhance visual partbased feature coherence and locality information. Our framework is novel in its ability to implicitly learn aligned semantics between modalities during the feature learning stage. The unified feature learning effectively utilizes textual data as a super-annotation signal for visual representation learning and automatically rejects irrelevant information. The entire AXM-Net is trained end-to-end on CUHK-PEDES data. We report results on two tasks, person search and cross-modal Re-ID. The AXM-Net outperforms the current state-of-theart (SOTA) methods and achieves 64.44% Rank@1 on the CUHK-PEDES test set. It also outperforms its competitors by >10% in cross-viewpoint text-to-image Re-ID scenarios on CrossRe-ID and CUHK-SYSU datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers22
- Towards Unified Text-based Person Retrieval: A Large-scale Multi-Attribute and Language Search BenchmarkShuyu Yang, Yinan Zhou, Zhedong Zheng, Yaxiong Wang et al.ACM MM 2023 · 162 citations
- An Empirical Study of CLIP for Text-Based Person SearchMin Cao, Yang Bai, Ziyin Zeng, Mang Ye et al.AAAI 2024 · 111 citations
- Adaptive Uncertainty-Based Learning for Text-Based Person RetrievalShenshen Li, Chen He, Xing Xu, Fumin Shen et al.AAAI 2024 · 59 citations
- Learning Comprehensive Representations with Richer Self for Text-to-Image Person Re-IdentificationShuanglin Yan, Neng Dong, Jun Liu, Liyan Zhang et al.ACM MM 2023 · 55 citations
- Unified Pre-training with Pseudo Texts for Text-To-Image Person Re-identificationZhiyin Shao, Xinyu Zhang, Changxing Ding, Jian Wang et al.ICCV 2023 · 46 citations
Builds on9
- Omni-Scale Feature Learning for Person Re-IdentificationKaiyang Zhou, Yongxin Yang, Andrea Cavallaro, Tao XiangICCV 2019 · 997 citations
- Mixed High-Order Attention Network for Person Re-IdentificationBinghui Chen, Weihong Deng, Jiani HuICCV 2019 · 392 citations
- Batch DropBlock Network for Person Re-Identification and BeyondZuozhuo Dai, Mingqiang Chen, Xiaodong Gu, Siyu Zhu et al.ICCV 2019 · 263 citations
- Second-Order Non-Local Attention Networks for Person Re-IdentificationBryan Bryan, Yuan Gong, Yizhe Zhang, Christian PoellabauerICCV 2019 · 196 citations
- Pose-Guided Multi-Granularity Attention Network for Text-Based Person SearchYa Jing, Chenyang Si, Junbo Wang, Wei Wang et al.AAAI 2020 · 182 citations
Related papers
- Cross-Modal Cross-Domain Moment Alignment Network for Person SearchYa Jing, Wei Wang, Liang Wang, Tieniu TanCVPR 2020
- DCEL: Deep Cross-modal Evidential Learning for Text-Based Person RetrievalShenshen Li, Xing Xu, Yang Yang, Fumin Shen et al.ACM MM 2023 · 56 citations
- Cross-modal Co-occurrence Attributes Alignments for Person Search by LanguageKai Niu, Linjiang Huang, Yan Huang, Peng Wang et al.ACM MM 2022 · 35 citations
- Multi-Modal Disordered Representation Learning Network for Description-Based Person SearchFan Yang, Wei Li, Menglong Yang, Binbin Liang et al.AAAI 2024 · 8 citations
- DSSL: Deep Surroundings-person Separation Learning for Text-based Person RetrievalAichun Zhu, Zijie Wang, Yifeng Li, Xili Wan et al.ACM MM 2021 · 274 citations
