AXM-Net: Implicit Cross-Modal Feature Alignment for Person Re-identification
Ammarah Farooq, Muhammad Awais, Josef Kittler, Syed Safwan Khalid
摘要
Cross-modal person re-identification (Re-ID) is critical for modern video surveillance systems. The key challenge is to align cross-modality representations induced by the semantic information present for a person and ignore background information. This work presents a novel convolutional neural network (CNN) based architecture designed to learn semantically aligned cross-modal visual and textual representations. The underlying building block, named AXM-Block, is a unified multi-layer network that dynamically exploits the multi-scale knowledge from both modalities and re-calibrates each modality according to shared semantics. To complement the convolutional design, contextual attention is applied in the text branch to manipulate long-term dependencies. Moreover, we propose a unique design to enhance visual partbased feature coherence and locality information. Our framework is novel in its ability to implicitly learn aligned semantics between modalities during the feature learning stage. The unified feature learning effectively utilizes textual data as a super-annotation signal for visual representation learning and automatically rejects irrelevant information. The entire AXM-Net is trained end-to-end on CUHK-PEDES data. We report results on two tasks, person search and cross-modal Re-ID. The AXM-Net outperforms the current state-of-theart (SOTA) methods and achieves 64.44% Rank@1 on the CUHK-PEDES test set. It also outperforms its competitors by >10% in cross-viewpoint text-to-image Re-ID scenarios on CrossRe-ID and CUHK-SYSU datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Towards Unified Text-based Person Retrieval: A Large-scale Multi-Attribute and Language Search BenchmarkShuyu Yang, Yinan Zhou, Zhedong Zheng, Yaxiong Wang 等ACM MM 2023 · 被引用 162 次
- An Empirical Study of CLIP for Text-Based Person SearchMin Cao, Yang Bai, Ziyin Zeng, Mang Ye 等AAAI 2024 · 被引用 111 次
- Adaptive Uncertainty-Based Learning for Text-Based Person RetrievalShenshen Li, Chen He, Xing Xu, Fumin Shen 等AAAI 2024 · 被引用 59 次
- Learning Comprehensive Representations with Richer Self for Text-to-Image Person Re-IdentificationShuanglin Yan, Neng Dong, Jun Liu, Liyan Zhang 等ACM MM 2023 · 被引用 55 次
- Unified Pre-training with Pseudo Texts for Text-To-Image Person Re-identificationZhiyin Shao, Xinyu Zhang, Changxing Ding, Jian Wang 等ICCV 2023 · 被引用 46 次
它引用的顶会 Paper9
- Omni-Scale Feature Learning for Person Re-IdentificationKaiyang Zhou, Yongxin Yang, Andrea Cavallaro, Tao XiangICCV 2019 · 被引用 997 次
- Mixed High-Order Attention Network for Person Re-IdentificationBinghui Chen, Weihong Deng, Jiani HuICCV 2019 · 被引用 392 次
- Batch DropBlock Network for Person Re-Identification and BeyondZuozhuo Dai, Mingqiang Chen, Xiaodong Gu, Siyu Zhu 等ICCV 2019 · 被引用 263 次
- Second-Order Non-Local Attention Networks for Person Re-IdentificationBryan Bryan, Yuan Gong, Yizhe Zhang, Christian PoellabauerICCV 2019 · 被引用 196 次
- Pose-Guided Multi-Granularity Attention Network for Text-Based Person SearchYa Jing, Chenyang Si, Junbo Wang, Wei Wang 等AAAI 2020 · 被引用 182 次
相关 Paper
- Cross-Modal Cross-Domain Moment Alignment Network for Person SearchYa Jing, Wei Wang, Liang Wang, Tieniu TanCVPR 2020
- DCEL: Deep Cross-modal Evidential Learning for Text-Based Person RetrievalShenshen Li, Xing Xu, Yang Yang, Fumin Shen 等ACM MM 2023 · 被引用 56 次
- Cross-modal Co-occurrence Attributes Alignments for Person Search by LanguageKai Niu, Linjiang Huang, Yan Huang, Peng Wang 等ACM MM 2022 · 被引用 35 次
- Multi-Modal Disordered Representation Learning Network for Description-Based Person SearchFan Yang, Wei Li, Menglong Yang, Binbin Liang 等AAAI 2024 · 被引用 8 次
- DSSL: Deep Surroundings-person Separation Learning for Text-based Person RetrievalAichun Zhu, Zijie Wang, Yifeng Li, Xili Wan 等ACM MM 2021 · 被引用 274 次
