Spatially Attentive Output Layer for Image Classification
Ildoo Kim, Woonhyuk Baek, Sungwoong Kim
Abstract
Most convolutional neural networks (CNNs) for image classification use a global average pooling (GAP) followed by a fully-connected (FC) layer for output logits. However, this spatial aggregation procedure inherently restricts the utilization of location-specific information at the output layer, although this spatial information can be beneficial for classification. In this paper, we propose a novel spatial output layer on top of the existing convolutional feature maps to explicitly exploit the location-specific output information. In specific, given the spatial feature maps, we replace the previous GAP-FC layer with a spatially attentive output layer (SAOL) by employing a attention mask on spatial logits. The proposed location-specific attention selectively aggregates spatial logits within a target region, which leads to not only the performance improvement but also spatially interpretable outputs. Moreover, the proposed SAOL also permits to fully exploit location-specific selfsupervision as well as self-distillation to enhance the generalization ability during training. The proposed SAOL with self-supervision and self-distillation can be easily plugged into existing CNNs. Experimental results on various classification tasks with representative architectures show consistent performance improvements by SAOL at almost the same computational cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Mind the Pad - CNNs Can Develop Blind SpotsBilal Alsallakh, Narine Kokhlikyan, Vivek Miglani, Jun Yuan et al.ICLR 2021 · 32 citations
- Self-supervised Spatial Reasoning on Multi-View Line DrawingsSiyuan Xiang, Anbang Yang, Yanfei Xue, Yaoqing Yang et al.CVPR 2022 · 4 citations
- Explore Visual Concept Formation for Image ClassificationShengzhou Xiong, Yihua Tan, Guoyou WangICML 2021 · 4 citations
- SVGformer: Representation Learning for Continuous Vector Graphics using TransformersDefu Cao, Zhaowen Wang, Jose Echevarria, Yan LiuCVPR 2023
Builds on3
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- Sharpen Focus: Learning With Attention Separability and ConsistencyLezi Wang, Ziyan Wu, Srikrishna Karanam, Kuan-Chuan Peng et al.ICCV 2019 · 37 citations
Related papers
- Distilling Global and Local Logits with Densely Connected RelationsYoumin Kim, Jinbae Park, Younho Jang, Muhammad Salman Ali et al.ICCV 2021 · 33 citations
- Spatial-Aware Token for Weakly Supervised Object LocalizationPingyu Wu, Wei Zhai, Yang Cao, Jiebo Luo et al.ICCV 2023 · 19 citations
- Towards Learning Spatially Discriminative Feature RepresentationsChaofei Wang, Jiayu Xiao, Yizeng Han, Qisen Yang et al.ICCV 2021 · 23 citations
- Residual Attention: A Simple but Effective Method for Multi-Label RecognitionKe Zhu, Jianxin WuICCV 2021 · 190 citations
- Slot Attention with Re-Initialization and Self-DistillationRongzhen Zhao, Yi Zhao, Juho Kannala, Joni PajarinenACM MM 2025 · 1 citation
