Category-aware Allocation Transformer for Weakly Supervised Object Localization
Zhiwei Chen, Jinren Ding, Liujuan Cao, Yunhang Shen, Shengchuan Zhang, Guannan Jiang, Rongrong Ji
Abstract
Weakly supervised object localization (WSOL) aims to localize objects based on only image-level labels as supervision. Recently, transformers have been introduced into WSOL, yielding impressive results. The self-attention mechanism and multilayer perceptron structure in transformers preserve long-range feature dependency, facilitating complete localization of the full object extent. However, current transformer-based methods predict bounding boxes using category-agnostic attention maps, which may lead to confused and noisy object localization. To address this issue, we propose a novel Category-aware Allocation TRansformer (CATR) that learns category-aware representations for specific objects and produces corresponding category-aware attention maps for object localization. First, we introduce a Category-aware Stimulation Module (CSM) to induce learnable category biases for self-attention maps, providing auxiliary supervision to guide the learning of more effective transformer representations. Second, we design an Object Constraint Module (OCM) to refine the object regions for the category-aware attention maps in a self-supervised manner. Extensive experiments on the CUB-200-2011 and ILSVRC datasets demonstrate that the proposed CATR achieves significant and consistent performance improvements over competing approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c961240d-df45-4a54-a76d-2f16625e2cd2Cited by top-tier papers3
- Occluded Person Re-identification via Saliency-Guided Patch TransferLei Tan, Jiaer Xia, Wenfeng Liu, Pingyang Dai et al.AAAI 2024 · 54 citations
- CAKE: Category Aware Knowledge Extraction for Open-Vocabulary Object DetectionShiyuan Ma, Donglin Qian, Kai Ye, Shengchuan ZhangAAAI 2025 · 8 citations
- Motion-Aware Caching for Efficient Autoregressive Video GenerationJing Xu, Yuexiao Ma, Xuzhe Zheng, WANG et al.ICML 2026
Builds on19
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- TransReID: Transformer-based Object Re-IdentificationShuting He, Hao Luo, Pichao Wang, Fan Wang et al.ICCV 2021 · 1,172 citations
- Conformer: Local Features Coupling Global Representations for Visual RecognitionZhiliang Peng, Wei Huang, Shanzhi Gu, Lingxi Xie et al.ICCV 2021 · 723 citations
Related papers
- LCTR: On Awakening the Local Continuity of Transformer for Weakly Supervised Object LocalizationZhiwei Chen, Changan Wang, Yabiao Wang, Guannan Jiang et al.AAAI 2022 · 61 citations
- TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object LocalizationWei Gao, Fang Wan, Xingjia Pan, Zhiliang Peng et al.ICCV 2021 · 260 citations
- Spatial-Aware Token for Weakly Supervised Object LocalizationPingyu Wu, Wei Zhai, Yang Cao, Jiebo Luo et al.ICCV 2023 · 19 citations
- Proxy Probing Decoder for Weakly Supervised Object Localization: A Baseline InvestigationJingyuan Xu, Hongtao Xie, Chuanbin Liu, Yongdong ZhangACM MM 2022 · 3 citations
- Foreground Activation Maps for Weakly Supervised Object LocalizationMeng Meng, Tianzhu Zhang, Qi Tian, Yongdong Zhang et al.ICCV 2021 · 65 citations
