TransHP: Image Classification with Hierarchical Prompting
Wenhao Wang, Yifan Sun, Wei Li, Yi Yang
Abstract
This paper explores a hierarchical prompting mechanism for the hierarchical image classification (HIC) task. Different from prior HIC methods, our hierarchical prompting is the first to explicitly inject ancestor-class information as a tokenized hint that benefits the descendant-class discrimination. We think it well imitates human visual recognition, i.e., humans may use the ancestor class as a prompt to draw focus on the subtle differences among descendant classes. We model this prompting mechanism into a Transformer with Hierarchical Prompting (TransHP). TransHP consists of three steps: 1) learning a set of prompt tokens to represent the coarse (ancestor) classes, 2) on-the-fly predicting the coarse class of the input image at an intermediate block, and 3) injecting the prompt token of the predicted coarse class into the intermediate feature. Though the parameters of TransHP maintain the same for all input images, the injected coarse-class prompt conditions (modifies) the subsequent feature extraction and encourages a dynamic focus on relatively subtle differences among the descendant classes. Extensive experiments show that TransHP improves image classification on accuracy (e.g., improving ViT-B/16 by +2.83% ImageNet classification accuracy), training data efficiency (e.g., +12.69% improvement under 10% ImageNet training data), and model explainability. Moreover, TransHP also performs favorably against prior HIC methods, showing that TransHP well exploits the hierarchical information. The code is available at: https://github.com/WangWenhao0716/TransHP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ff2d9b6c-ebbf-453a-84c4-05e268d008e9Cited by top-tier papers10
- ChatGPT-Powered Hierarchical Comparisons for Image ClassificationZhiyuan Ren, Yiyang Su, Xiaoming LiuNeurIPS 2023 · 54 citations
- Hybrid Token Compression for Vision-Language ModelsJusheng Zhang, Xiaoyang Guo, Kaitong Cai, Qinhan Lv et al.CVPR 2026 · 23 citations
- Bayesian-guided Label Mapping for Visual ReprogrammingChengyi Cai, Zesheng Ye, Lei Feng, Jianzhong Qi et al.NeurIPS 2024 · 14 citations
- InsVP: Efficient Instance Visual Prompting from Image ItselfZichen Liu, Yuxin Peng, Jiahuan ZhouACM MM 2024 · 5 citations
- Bidirectional Logits Tree: Pursuing Granularity Reconcilement in Fine-Grained ClassificationZhiguang Lu, Qianqian Xu, Shilong Bao, Zhiyong Yang et al.AAAI 2025 · 1 citation
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick et al.ICLR 2022 · 1,182 citations
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang et al.CVPR 2022 · 635 citations
Related papers
- Scalable Vision Transformers with Hierarchical PoolingZizheng Pan, Bohan Zhuang, Jing Liu, Haoyu He et al.ICCV 2021 · 154 citations
- Learning Expressive Prompting With Residuals for Vision TransformersRajshekhar Das, Yonatan Dukler, Avinash Ravichandran, Ashwin SwaminathanCVPR 2023
- Hierarchical Prompt Learning for Multi-Task LearningYajing Liu, Yuning Lu, Hao Liu, Yaozu An et al.CVPR 2023
- Token Coordinated Prompt Attention is Needed for Visual PromptingZichen Liu, Xu Zou, Gang Hua, Jiahuan ZhouICML 2025
- Analyzing Vision Transformers for Image Classification in Class Embedding SpaceMartina G. Vilas, Timothy Schaumlöffel, Gemma RoigNeurIPS 2023 · 43 citations
