Inferring Prototypes for Multi-Label Few-Shot Image Classification with Word Vector Guided Attention
Kun Yan, Chenbin Zhang, Jun Hou, Ping Wang, Zied Bouraoui, Shoaib Jameel, Steven Schockaert
Abstract
Multi-label few-shot image classification (ML-FSIC) is the task of assigning descriptive labels to previously unseen images, based on a small number of training examples. A key feature of the multi-label setting is that images often have multiple labels, which typically refer to different regions of the image. When estimating prototypes, in a metric-based setting, it is thus important to determine which regions are relevant for which labels, but the limited amount of training data makes this highly challenging. As a solution, in this paper, we propose to use word embeddings as a form of prior knowledge about the meaning of the labels. In particular, visual prototypes are obtained by aggregating the local feature maps of the support images, using an attention mechanism that relies on the label embeddings. As an important advantage, our model can infer prototypes for unseen labels without the need for fine-tuning any model parameters, which demonstrates its strong generalization abilities. Experiments on COCO and PASCAL VOC furthermore show that our model substantially improves the current state-of-the-art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0f91b54d-47a2-499c-ac36-9b35062cd713Cited by top-tier papers6
- Making Large Vision Language Models to Be Good Few-Shot LearnersFan Liu, Wenwen Cai, Jian Huo, Chuanyi Zhang et al.AAAI 2025 · 7 citations
- Multi-Label Few-Shot Image Classification via Pairwise Feature Augmentation and Flexible Prompt LearningHan Liu, Yuanyuan Wang, Xiaotong Zhang, Feng Zhang et al.AAAI 2025 · 3 citations
- Distilling Semantic Concept Embeddings from Contrastively Fine-Tuned Language ModelsNa Li, Hanane Kteich, Zied Bouraoui, Steven SchockaertSIGIR 2023 · 3 citations
- Semantic Prompt for Few-Shot Image RecognitionCVPR 2023
- Rethinking BCE Loss for Multi-Label Image Recognition with Fine-TuningAo Zhou, Zhiwei Jiang, Zifeng Cheng, Cong Wang et al.CVPR 2026
Builds on8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Cross-Modality Attention with Semantic Graph Embedding for Multi-Label ClassificationRenchun You, Zhiyao Guo, Lei Cui, Xiang Long et al.AAAI 2020 · 221 citations
- Multi-Label Classification with Label Graph SuperimposingYa Wang, Dongliang He, Fu Li, Xiang Long et al.AAAI 2020 · 192 citations
- General Multi-Label Image Classification With TransformersJack Lanchantin, Tianlu Wang, Vicente Ordonez, Yanjun QiCVPR 2021
- Pareto Self-Supervised Training for Few-Shot LearningZhengyu Chen, Jixie Ge, Heshen Zhan, Siteng Huang et al.CVPR 2021
Related papers
- MIANet: Aggregating Unbiased Instance and General Information for Few-Shot Semantic SegmentationYong Yang, Qiong Chen, Yuan Feng, Tianlin HuangCVPR 2023
- Discovering Human Interactions With Novel Objects via Zero-Shot LearningSuchen Wang, Kim-Hui Yap, Junsong Yuan, Yap-Peng TanCVPR 2020
- Epsilon: Exploring Comprehensive Visual-Semantic Projection for Multi-Label Zero-Shot LearningZiming Liu, Jingcai Guo, Song Guo, Xiaocheng LuAAAI 2025 · 6 citations
- Semantic-Aware Knowledge Distillation for Few-Shot Class-Incremental LearningAli Cheraghian, Shafin Rahman, Pengfei Fang, Soumava Kumar Roy et al.CVPR 2021
- Breaking Immutable: Information-Coupled Prototype Elaboration for Few-Shot Object DetectionXiaonan Lu, Wenhui Diao, Yongqiang Mao, Junxi Li et al.AAAI 2023 · 66 citations
