Category-aware Allocation Transformer for Weakly Supervised Object Localization
Zhiwei Chen, Jinren Ding, Liujuan Cao, Yunhang Shen, Shengchuan Zhang, Guannan Jiang, Rongrong Ji
摘要
Weakly supervised object localization (WSOL) aims to localize objects based on only image-level labels as supervision. Recently, transformers have been introduced into WSOL, yielding impressive results. The self-attention mechanism and multilayer perceptron structure in transformers preserve long-range feature dependency, facilitating complete localization of the full object extent. However, current transformer-based methods predict bounding boxes using category-agnostic attention maps, which may lead to confused and noisy object localization. To address this issue, we propose a novel Category-aware Allocation TRansformer (CATR) that learns category-aware representations for specific objects and produces corresponding category-aware attention maps for object localization. First, we introduce a Category-aware Stimulation Module (CSM) to induce learnable category biases for self-attention maps, providing auxiliary supervision to guide the learning of more effective transformer representations. Second, we design an Object Constraint Module (OCM) to refine the object regions for the category-aware attention maps in a self-supervised manner. Extensive experiments on the CUB-200-2011 and ILSVRC datasets demonstrate that the proposed CATR achieves significant and consistent performance improvements over competing approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Occluded Person Re-identification via Saliency-Guided Patch TransferLei Tan, Jiaer Xia, Wenfeng Liu, Pingyang Dai 等AAAI 2024 · 被引用 54 次
- CAKE: Category Aware Knowledge Extraction for Open-Vocabulary Object DetectionShiyuan Ma, Donglin Qian, Kai Ye, Shengchuan ZhangAAAI 2025 · 被引用 8 次
- Motion-Aware Caching for Efficient Autoregressive Video GenerationJing Xu, Yuexiao Ma, Xuzhe Zheng, WANG 等ICML 2026
它引用的顶会 Paper19
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 被引用 2,072 次
- TransReID: Transformer-based Object Re-IdentificationShuting He, Hao Luo, Pichao Wang, Fan Wang 等ICCV 2021 · 被引用 1,172 次
- Conformer: Local Features Coupling Global Representations for Visual RecognitionZhiliang Peng, Wei Huang, Shanzhi Gu, Lingxi Xie 等ICCV 2021 · 被引用 723 次
相关 Paper
- LCTR: On Awakening the Local Continuity of Transformer for Weakly Supervised Object LocalizationZhiwei Chen, Changan Wang, Yabiao Wang, Guannan Jiang 等AAAI 2022 · 被引用 61 次
- TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object LocalizationWei Gao, Fang Wan, Xingjia Pan, Zhiliang Peng 等ICCV 2021 · 被引用 260 次
- Spatial-Aware Token for Weakly Supervised Object LocalizationPingyu Wu, Wei Zhai, Yang Cao, Jiebo Luo 等ICCV 2023 · 被引用 19 次
- Proxy Probing Decoder for Weakly Supervised Object Localization: A Baseline InvestigationJingyuan Xu, Hongtao Xie, Chuanbin Liu, Yongdong ZhangACM MM 2022 · 被引用 3 次
- Foreground Activation Maps for Weakly Supervised Object LocalizationMeng Meng, Tianzhu Zhang, Qi Tian, Yongdong Zhang 等ICCV 2021 · 被引用 65 次
