TransMatcher: Deep Image Matching Through Transformers for Generalizable Person Re-identification
Shengcai Liao, Ling Shao
摘要
Transformers have recently gained increasing attention in computer vision. However, existing studies mostly use Transformers for feature representation learning, e.g. for image classification and dense predictions, and the generalizability of Transformers is unknown. In this work, we further investigate the possibility of applying Transformers for image matching and metric learning given pairs of images. We find that the Vision Transformer (ViT) and the vanilla Transformer with decoders are not adequate for image matching due to their lack of image-to-image attention. Thus, we further design two naive solutions, i.e. query-gallery concatenation in ViT, and query-gallery cross-attention in the vanilla Transformer. The latter improves the performance, but it is still limited. This implies that the attention mechanism in Transformers is primarily designed for global feature aggregation, which is not naturally suitable for image matching. Accordingly, we propose a new simplified decoder, which drops the full attention implementation with the softmax weighting, keeping only the query-key similarity computation. Additionally, global max pooling and a multilayer perceptron (MLP) head are applied to decode the matching result. This way, the simplified decoder is computationally more efficient, while at the same time more effective for image matching. The proposed method, called TransMatcher, achieves state-of-the-art performance in generalizable person re-identification, with up to 6.1% and 5.7% performance gains in Rank-1 and mAP, respectively, on several popular datasets. Code is available at https://github.com/ShengcaiLiao/QAConv .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Graph Sampling Based Deep Metric Learning for Generalizable Person Re-IdentificationShengcai Liao, Ling ShaoCVPR 2022 · 被引用 122 次
- PLIP: Language-Image Pre-training for Person Representation LearningJialong Zuo, Jiahao Hong, Feng Zhang, Changqian Yu 等NeurIPS 2024 · 被引用 96 次
- Part-Aware Transformer for Generalizable Person Re-identificationHao Ni, Yuke Li, Lianli Gao, Heng Tao Shen 等ICCV 2023 · 被引用 87 次
- Cloning Outfits from Real-World Images to 3D Characters for Generalizable Person Re-IdentificationYanan Wang, Xuezhi Liang, Shengcai LiaoCVPR 2022 · 被引用 35 次
- Generalizable Person Re-identification via Balancing Alignment and UniformityYoonki Cho, Jaeyoon Kim, Woo Jae Kim, Junsik Jung 等NeurIPS 2024 · 被引用 21 次
它引用的顶会 Paper10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- Omni-Scale Feature Learning for Person Re-IdentificationKaiyang Zhou, Yongxin Yang, Andrea Cavallaro, Tao XiangICCV 2019 · 被引用 997 次
相关 Paper
- Compact Transformer Tracker with Correlative Masked ModelingZikai Song, Run Luo, Junqing Yu, Yi-Ping Phoebe Chen 等AAAI 2023 · 被引用 136 次
- TransMix: Attend to Mix for Vision TransformersJieneng Chen, Shuyang Sun, Ju He, Philip H. S. Torr 等CVPR 2022
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski 等AAAI 2022 · 被引用 529 次
- TransforMatcher: Match-to-Match Attention for Semantic CorrespondenceSeungwook Kim, Juhong Min, Minsu ChoCVPR 2022 · 被引用 26 次
- Rethinking Spatial Dimensions of Vision TransformersByeongho Heo, Sangdoo Yun, Dongyoon Han, Sanghyuk Chun 等ICCV 2021 · 被引用 733 次
