Compressed Self-Attention for Deep Metric Learning
Ziye Chen, Mingming Gong, Yanwu Xu, Chaohui Wang, Kun Zhang, Bo Du
Abstract
In this paper, we aim to enhance self-attention (SA) mechanism for deep metric learning in visual perception, by capturing richer contextual dependencies in visual data. To this end, we propose a novel module, named compressed self-attention (CSA), which significantly reduces the computation and memory cost with a neglectable decrease in accuracy with respect to the original SA mechanism, thanks to the following two characteristics: i) it only needs to compute a small number of base attention maps for a small number of base feature vectors; and ii) the output at each spatial location can be simply obtained by an adaptive weighted average of the outputs calculated from the base attention maps. The high computational efficiency of CSA enables the application to high-resolution shallow layers in convolutional neural networks with little additional cost. In addition, CSA makes it practical to further partition the feature maps into groups along the channel dimension and compute attention maps for features in each group separately, thus increasing the diversity of long-range dependencies and accordingly boosting the accuracy. We evaluate the performance of CSA via extensive experiments on two metric learning tasks: person re-identification and local descriptor learning. Qualitative and quantitative comparisons with latest methods demonstrate the significance of CSA in this topic.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- Expectation-Maximization Attention Networks for Semantic SegmentationXia Li, Zhisheng Zhong, Jianlong Wu, Yibo Yang et al.ICCV 2019 · 639 citations
- Explicitly Modeled Attention Maps for Image ClassificationAndong Tan, Duc Tam Nguyen, Maximilian Dax, Matthias Nießner et al.AAAI 2021 · 10 citations
- Learning Deep Local Features with Multiple Dynamic Attentions for Large-Scale Image RetrievalHui Wu, Min Wang, Wengang Zhou, Houqiang LiICCV 2021 · 26 citations
- Self-Critical Attention Learning for Person Re-IdentificationGuangyi Chen, Chunze Lin, Liangliang Ren, Jiwen Lu et al.ICCV 2019 · 146 citations
- Core Context Aware Transformers for Long Context Language ModelingYaofo Chen, Zeng You, Shuhai Zhang, Haokun Li et al.ICML 2025
