Less Attention is More: Prompt Transformer for Generalized Category Discovery
Wei Zhang, Baopeng Zhang, Zhu Teng, Wenxin Luo, Junnan Zou, Jianping Fan
Abstract
Generalized Category Discovery (GCD) typically relies on the pre-trained Vision Transformer (ViT) to extract features from a global receptive field, followed by contrastive learning to simultaneously classify unlabeled known classes and unknown classes without priors. Owing to the deficiency in the modeling capacity for inner-patch local information within ViT, current methods primarily focus on discriminative features at global level. This results in a model with more yet scattered attention, where neither excessive nor insufficient focus can grasp subtle differences to classify fine-grained unknown and known categories. To address this issue, we propose the AptGCD to deliver apt attention for GCD. It mimics the human brain how leveraging visual perception to refine local attention and comprehend global context by proposing a Meta Visual Prompt (MVP) and Prompt Transformer (PT). MVP is introduced into GCD for the first time, refining channel-level attention, while adaptively self-learning unique inner-patch features as prompts to achieve local visual modeling for our prompt transformer. Yet, relying solely on detailed features can lead to skewed judgments. Hence, PT harmonizes local and global representations, guiding the model's interpretation of features through broader contexts, thereby capturing more useful details with less attention. Extensive experiments on seven datasets demonstrate that AptGCD outperforms current methods, it achieves an average proportional 'New' accuracy improvement of approximately 9.2% over SOTA method on the all four fine-grained datasets, establishing a new standard in the field. The code is available at https://github.com/wendy26zhang/AptGCD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 26165261-cae7-4d43-9933-9d1010c30a42Cited by top-tier papers4
- A Hidden Stumbling Block in Generalized Category Discovery: Distracted AttentionQiyu Xu, Zhanxuan Hu, Yu Duan, Ercheng Pei et al.ICCV 2025 · 5 citations
- SpectralGCD: Spectral Concept Selection and Cross-modal Representation Learning for Generalized Category DiscoveryLorenzo Caselli, Marco Mistretta, Simone Magistri, Andrew D. BagdanovICLR 2026 · 3 citations
- The Devil Is in Gradient Entanglement: Energy-Aware Gradient Coordinator for Robust Generalized Category DiscoveryHaiyang Zheng, Nan Pu, Yaqi Cai, Teng Long et al.CVPR 2026 · 1 citation
- The Finer the Better: Towards Granular-aware Open-set Domain GeneralizationYunyun Wang, Zheng Duan, Xinyue Liao, Ke-Jia Chen et al.AAAI 2026
Builds on22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- Open-Set Recognition: A Good Closed-Set Classifier is All You NeedSagar Vaze, Kai Han, Andrea Vedaldi, Andrew ZissermanICLR 2022 · 594 citations
- Conditional Positional Encodings for Vision TransformersXiangxiang Chu, Zhi Tian, Bo Zhang, Xinlong Wang et al.ICLR 2023 · 406 citations
Related papers
- TAR: Token-Aware Refinement for Fine-grained Generalized Category DiscoveryXingyu Yang, Yu Zhang, Siya Mi, Xiu-Shen WeiCVPR 2026
- Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play EnhancementQiyuan Dai, Hanzhuo Huang, Yu Wu, Sibei YangCVPR 2025
- PromptCAL: Contrastive Affinity Learning via Auxiliary Prompts for Generalized Novel Category DiscoverySheng Zhang, Salman H. Khan, Zhiqiang Shen, Muzammal Naseer et al.CVPR 2023
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski et al.AAAI 2022 · 529 citations
- SPTNet: An Efficient Alternative Framework for Generalized Category Discovery with Spatial Prompt TuningHongjun Wang, Sagar Vaze, Kai HanICLR 2024 · 57 citations
