Proxy Probing Decoder for Weakly Supervised Object Localization: A Baseline Investigation
Jingyuan Xu, Hongtao Xie, Chuanbin Liu, Yongdong Zhang
Abstract
Weakly supervised object localization (WSOL) aims to localize the object with only image category labels. Existing methods generally fine-tune the models with manually selected training epochs and subjective loss functions to mitigate the partial activation problem of the classification-based model. However, such fine-tuning scheme would cause the model to degrade, e.g. affect the classification performance and generalization capabilities of the pre-trained model. In this paper, we propose a novel method named Proxy Probing Decoder (PPD) to meet these challenges, which utilizes the segmentation property of self-attention map in the self-supervised vision transformer and breaks through model fine-tuning with a novel proxy probing decoder. Specifically, we utilize the self-supervised vision transformer to capture long-range dependencies and avoid partial activation. Then we simply adopt a proxy consisting of a series of decoding layers to transform the feature representations into the heatmap of the objects' foreground and conduct localization. The backbone parameters are frozen during training while the proxy is used to decode the feature and localize the object. In this way, the vision transformer model can maintain the feature representation capabilities and only the proxy is required for adapting to the task. Without bells and whistles, our framework achieves 55.0% Top-1 Loc on the ILSVRC2012 dataset and 78.8% Top-1 Loc on the CUB-200-2011 dataset, which surpasses state-of-the-art by a large margin and provides a simple baseline. Codes and models will be available on Github.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 40df7c32-880c-4420-b0c2-5c1601a36867Cited by top-tier papers3
- WeakSAM: Segment Anything Meets Weakly-supervised Instance-level RecognitionLianghui Zhu, Junwei Zhou, Yan Liu, Xin Hao et al.ACM MM 2024 · 21 citations
- Rethinking the Localization in Weakly Supervised Object LocalizationRui Xu, Yong Luo, Han Hu, Bo Du et al.ACM MM 2023 · 8 citations
- FDCNet: Feature Drift Compensation Network for Class-Incremental Weakly Supervised Object LocalizationSejin Park, Taehyung Lee, Yeejin Lee, Byeongkeun KangACM MM 2023 · 3 citations
Related papers
- TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object LocalizationWei Gao, Fang Wan, Xingjia Pan, Zhiliang Peng et al.ICCV 2021 · 260 citations
- Category-aware Allocation Transformer for Weakly Supervised Object LocalizationZhiwei Chen, Jinren Ding, Liujuan Cao, Yunhang Shen et al.ICCV 2023 · 15 citations
- Spatial-Aware Token for Weakly Supervised Object LocalizationPingyu Wu, Wei Zhai, Yang Cao, Jiebo Luo et al.ICCV 2023 · 19 citations
- TriLite: Efficient Weakly Supervised Object Localization with Universal Visual Features and Tri-Region DisentanglementArian Sabaghi, José OramasCVPR 2026
- LCTR: On Awakening the Local Continuity of Transformer for Weakly Supervised Object LocalizationZhiwei Chen, Changan Wang, Yabiao Wang, Guannan Jiang et al.AAAI 2022 · 61 citations
