Generative Prompt Model for Weakly Supervised Object Localization
Yuzhong Zhao, Qixiang Ye, Weijia Wu, Chunhua Shen, Fang Wan
摘要
Weakly supervised object localization (WSOL) remains challenging when learning object localization models from image category labels. Conventional methods that discriminatively train activation models ignore representative yet less discriminative object parts. In this study, we propose a generative prompt model (GenPromp), defining the first generative pipeline to localize less discriminative object parts by formulating WSOL as a conditional image denoising procedure. During training, GenPromp converts image category labels to learnable prompt embeddings which are fed to a generative model to conditionally recover the input image with noise and learn representative embeddings. During inference, GenPromp combines the representative embeddings with discriminative embeddings (queried from an off-the-shelf vision-language model) for both representative and discriminative capacity. The combined embeddings are finally used to generate multi-scale high-quality attention maps, which facilitate localizing full object extent. Experiments on CUB-200-2011 and ILSVRC show that GenPromp respectively outperforms the best discriminative models by 5.2% and 5.6% (Top-1 Loc), setting a solid baseline for WSOL with the generative model. Code is available at https://github.com/callsys/GenPromp .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- DiffuMask: Synthesizing Images with Pixel-level Annotations for Semantic Segmentation Using Diffusion ModelsWeijia Wu, Yuzhong Zhao, Mike Zheng Shou, Hong Zhou 等ICCV 2023 · 被引用 198 次
- DatasetDM: Synthesizing Data with Perception Annotations Using Diffusion ModelsWeijia Wu, Yuzhong Zhao, Hao Chen, Yuchao Gu 等NeurIPS 2023 · 被引用 191 次
- Keep Various Trajectories: Promoting Exploration of Ensemble Policies in Continuous ControlChao Li, Chen Gong, Qiang He, Xinwen HouNeurIPS 2023 · 被引用 8 次
- Leveraging Prior Knowledge of Diffusion Model for Person SearchGiyeol Kim, Sooyoung Yang, Jihyong Oh, Myungjoo Kang 等ICCV 2025 · 被引用 2 次
- Selective Contrastive Learning for Weakly Supervised Affordance GroundingWonJun Moon, Hyun Seok Seong, Jae-Pil HeoICCV 2025 · 被引用 1 次
它引用的顶会 Paper34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Foreground Activation Maps for Weakly Supervised Object LocalizationMeng Meng, Tianzhu Zhang, Qi Tian, Yongdong Zhang 等ICCV 2021 · 被引用 65 次
- Online Refinement of Low-level Feature Based Activation Map for Weakly Supervised Object LocalizationJinheng Xie, Cheng Luo, Xiangping Zhu, Ziqi Jin 等ICCV 2021 · 被引用 61 次
- Erasing Integrated Learning: A Simple Yet Effective Approach for Weakly Supervised Object LocalizationJinjie Mai, Meng Yang, Wenfeng LuoCVPR 2020
- TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object LocalizationWei Gao, Fang Wan, Xingjia Pan, Zhiliang Peng 等ICCV 2021 · 被引用 260 次
- CoPL: Contextual Prompt Learning for Vision-Language UnderstandingKoustava Goswami, Srikrishna Karanam, Prateksha Udhayanan, K. J. Joseph 等AAAI 2024 · 被引用 20 次
