Less Attention is More: Prompt Transformer for Generalized Category Discovery
Wei Zhang, Baopeng Zhang, Zhu Teng, Wenxin Luo, Junnan Zou, Jianping Fan
摘要
Generalized Category Discovery (GCD) typically relies on the pre-trained Vision Transformer (ViT) to extract features from a global receptive field, followed by contrastive learning to simultaneously classify unlabeled known classes and unknown classes without priors. Owing to the deficiency in the modeling capacity for inner-patch local information within ViT, current methods primarily focus on discriminative features at global level. This results in a model with more yet scattered attention, where neither excessive nor insufficient focus can grasp subtle differences to classify fine-grained unknown and known categories. To address this issue, we propose the AptGCD to deliver apt attention for GCD. It mimics the human brain how leveraging visual perception to refine local attention and comprehend global context by proposing a Meta Visual Prompt (MVP) and Prompt Transformer (PT). MVP is introduced into GCD for the first time, refining channel-level attention, while adaptively self-learning unique inner-patch features as prompts to achieve local visual modeling for our prompt transformer. Yet, relying solely on detailed features can lead to skewed judgments. Hence, PT harmonizes local and global representations, guiding the model's interpretation of features through broader contexts, thereby capturing more useful details with less attention. Extensive experiments on seven datasets demonstrate that AptGCD outperforms current methods, it achieves an average proportional 'New' accuracy improvement of approximately 9.2% over SOTA method on the all four fine-grained datasets, establishing a new standard in the field. The code is available at https://github.com/wendy26zhang/AptGCD .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- A Hidden Stumbling Block in Generalized Category Discovery: Distracted AttentionQiyu Xu, Zhanxuan Hu, Yu Duan, Ercheng Pei 等ICCV 2025 · 被引用 5 次
- SpectralGCD: Spectral Concept Selection and Cross-modal Representation Learning for Generalized Category DiscoveryLorenzo Caselli, Marco Mistretta, Simone Magistri, Andrew D. BagdanovICLR 2026 · 被引用 3 次
- The Devil Is in Gradient Entanglement: Energy-Aware Gradient Coordinator for Robust Generalized Category DiscoveryHaiyang Zheng, Nan Pu, Yaqi Cai, Teng Long 等CVPR 2026 · 被引用 1 次
- The Finer the Better: Towards Granular-aware Open-set Domain GeneralizationYunyun Wang, Zheng Duan, Xinyue Liao, Ke-Jia Chen 等AAAI 2026
它引用的顶会 Paper22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
- Open-Set Recognition: A Good Closed-Set Classifier is All You NeedSagar Vaze, Kai Han, Andrea Vedaldi, Andrew ZissermanICLR 2022 · 被引用 594 次
- Conditional Positional Encodings for Vision TransformersXiangxiang Chu, Zhi Tian, Bo Zhang, Xinlong Wang 等ICLR 2023 · 被引用 406 次
相关 Paper
- TAR: Token-Aware Refinement for Fine-grained Generalized Category DiscoveryXingyu Yang, Yu Zhang, Siya Mi, Xiu-Shen WeiCVPR 2026
- Adaptive Part Learning for Fine-Grained Generalized Category Discovery: A Plug-and-Play EnhancementQiyuan Dai, Hanzhuo Huang, Yu Wu, Sibei YangCVPR 2025
- PromptCAL: Contrastive Affinity Learning via Auxiliary Prompts for Generalized Novel Category DiscoverySheng Zhang, Salman H. Khan, Zhiqiang Shen, Muzammal Naseer 等CVPR 2023
- TransFG: A Transformer Architecture for Fine-Grained RecognitionJu He, Jieneng Chen, Shuai Liu, Adam Kortylewski 等AAAI 2022 · 被引用 529 次
- SPTNet: An Efficient Alternative Framework for Generalized Category Discovery with Spatial Prompt TuningHongjun Wang, Sagar Vaze, Kai HanICLR 2024 · 被引用 57 次
