Coarse2Fine: Local Consistency Aware Re-prediction for Weakly Supervised Object Localization
Yixuan Pan, Yao Yao, Yichao Cao, Chongjin Chen, Xiaobo Lu
Abstract
Weakly supervised object localization aims to localize objects of interest by using only image-level labels. Existing methods generally segment the activation map by threshold to obtain mask and generate a bounding box. However, the activation map is locally inconsistent, i.e., similar neighboring pixels of the same object are not equally activated, which leads to the blurred boundary issue: the localization result is sensitive to the threshold, and the mask obtained directly from the activation map loses the fine contours of the object, making it difficult to obtain a tight bounding box. In this paper, we introduce the Local Consistency Aware Re-prediction (LCAR) framework, which aims to recover the complete fine object mask from the locally inconsistent activation map and hence obtain a tight bounding box. To this end, we propose the self-guided re-prediction module (SGRM), which employs a novel Aggregation Net with dynamic weights to replace the post-processing of threshold segmentation. To derive more reliable pseudo labels from the activation map to supervise the SGRM, we further design an affinity refinement module (ARM) that utilizes the original image feature to better align the activation map with the image contents and design a selfdistillation CAM (SD-CAM) to alleviate the localizer dependence on saliency. Experiments demonstrate that our LCAR outperforms the state-of-the-art on both the CUB-200-2011 and ILSVRC datasets, achieving 95.9% and 70.7% of GT-Know localization accuracy, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 495477b6-2cf8-4fdd-a3e5-62fc386f3ed9Builds on15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object LocalizationWei Gao, Fang Wan, Xingjia Pan, Zhiliang Peng et al.ICCV 2021 · 260 citations
- Learning Affinity from Attention: End-to-End Weakly-Supervised Semantic Segmentation with TransformersLixiang Ru, Yibing Zhan, Baosheng Yu, Bo DuCVPR 2022 · 257 citations
Related papers
- CREAM: Weakly Supervised Object Localization via Class RE-Activation MappingJilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng et al.CVPR 2022 · 38 citations
- Bridging the Gap between Classification and Localization for Weakly Supervised Object LocalizationEunji Kim, Siwon Kim, Jungbeom Lee, Hyunwoo Kim et al.CVPR 2022 · 44 citations
- Online Refinement of Low-level Feature Based Activation Map for Weakly Supervised Object LocalizationJinheng Xie, Cheng Luo, Xiangping Zhu, Ziqi Jin et al.ICCV 2021 · 61 citations
- Self-Supervised Object Localization with Joint Graph PartitionYukun Su, Guosheng Lin, Yun Hao, Yiwen Cao et al.AAAI 2022 · 17 citations
- Category-aware Allocation Transformer for Weakly Supervised Object LocalizationZhiwei Chen, Jinren Ding, Liujuan Cao, Yunhang Shen et al.ICCV 2023 · 15 citations
