Spatially Attentive Output Layer for Image Classification
Ildoo Kim, Woonhyuk Baek, Sungwoong Kim
摘要
Most convolutional neural networks (CNNs) for image classification use a global average pooling (GAP) followed by a fully-connected (FC) layer for output logits. However, this spatial aggregation procedure inherently restricts the utilization of location-specific information at the output layer, although this spatial information can be beneficial for classification. In this paper, we propose a novel spatial output layer on top of the existing convolutional feature maps to explicitly exploit the location-specific output information. In specific, given the spatial feature maps, we replace the previous GAP-FC layer with a spatially attentive output layer (SAOL) by employing a attention mask on spatial logits. The proposed location-specific attention selectively aggregates spatial logits within a target region, which leads to not only the performance improvement but also spatially interpretable outputs. Moreover, the proposed SAOL also permits to fully exploit location-specific selfsupervision as well as self-distillation to enhance the generalization ability during training. The proposed SAOL with self-supervision and self-distillation can be easily plugged into existing CNNs. Experimental results on various classification tasks with representative architectures show consistent performance improvements by SAOL at almost the same computational cost.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Mind the Pad - CNNs Can Develop Blind SpotsBilal Alsallakh, Narine Kokhlikyan, Vivek Miglani, Jun Yuan 等ICLR 2021 · 被引用 32 次
- Self-supervised Spatial Reasoning on Multi-View Line DrawingsSiyuan Xiang, Anbang Yang, Yanfei Xue, Yaoqing Yang 等CVPR 2022 · 被引用 4 次
- Explore Visual Concept Formation for Image ClassificationShengzhou Xiong, Yihua Tan, Guoyou WangICML 2021 · 被引用 4 次
- SVGformer: Representation Learning for Continuous Vector Graphics using TransformersDefu Cao, Zhaowen Wang, Jose Echevarria, Yan LiuCVPR 2023
它引用的顶会 Paper3
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
- Sharpen Focus: Learning With Attention Separability and ConsistencyLezi Wang, Ziyan Wu, Srikrishna Karanam, Kuan-Chuan Peng 等ICCV 2019 · 被引用 37 次
相关 Paper
- Distilling Global and Local Logits with Densely Connected RelationsYoumin Kim, Jinbae Park, Younho Jang, Muhammad Salman Ali 等ICCV 2021 · 被引用 33 次
- Spatial-Aware Token for Weakly Supervised Object LocalizationPingyu Wu, Wei Zhai, Yang Cao, Jiebo Luo 等ICCV 2023 · 被引用 19 次
- Towards Learning Spatially Discriminative Feature RepresentationsChaofei Wang, Jiayu Xiao, Yizeng Han, Qisen Yang 等ICCV 2021 · 被引用 23 次
- Residual Attention: A Simple but Effective Method for Multi-Label RecognitionKe Zhu, Jianxin WuICCV 2021 · 被引用 190 次
- Slot Attention with Re-Initialization and Self-DistillationRongzhen Zhao, Yi Zhao, Juho Kannala, Joni PajarinenACM MM 2025 · 被引用 1 次
