MPA: Multimodal Prototype Augmentation for Few-Shot Learning
Liwen Wu, Wei Wang, Lei Zhao, Zhan Gao, Qika Lin, Shaowen Yao, Zuozhu Liu, Bin Pu
Abstract
Recently, Few-shot Learning (FSL) has become a popular task that aims to recognize new classes from only a few labeled examples and has been widely applied in fields such as natural science, remote sensing, and medical images. However, most existing methods focus only on the visual modality and compute prototypes directly from raw support images, which lack comprehensive and rich multimodal information. To address these limitations, we propose a novel Multimodal Prototype Augmentation FSL framework called MPA, including LLM-based Multi-Variant Semantic Enhancement (LMSE), Hierarchical Multi-View Augmentation (HMA), and an Adaptive Uncertain Class Absorber (AUCA). LMSE leverages large language models to generate diverse paraphrased category descriptions, enriching the support set with additional semantic cues. HMA exploits both natural and multi-view augmentations to enhance feature diversity (e.g., changes in viewing distance, camera angles, and lighting conditions). AUCA models uncertainty by introducing uncertain classes via interpolation and Gaussian sampling, effectively absorbing uncertain samples. Extensive experiments on four single-domain and six cross-domain FSL benchmarks demonstrate that MPA achieves superior performance compared to existing state-of-the-art methods across most settings. Notably, MPA surpasses the second-best method by 12.29% and 24.56% in the single-domain and cross-domain setting, respectively, in the 5-way 1-shot setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 35ca267e-8a8b-4645-a58a-cb89e7467aacBuilds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Ranking Distance Calibration for Cross-Domain Few-Shot LearningPan Li, Shaogang Gong, Chengjie Wang, Yanwei FuCVPR 2022 · 59 citations
- Attribute Surrogates Learning and Spectral Tokens Pooling in Transformers for Few-shot LearningYangji He, Weihan Liang, Dongyang Zhao, Hong-Yu Zhou et al.CVPR 2022 · 58 citations
- Class-Aware Patch Embedding Adaptation for Few-Shot Image ClassificationFusheng Hao, Fengxiang He, Liu Liu, Fuxiang Wu et al.ICCV 2023 · 56 citations
- Context-Aware Meta-LearningChristopher Fifty, Dennis Duan, Ronald G. Junkins, Ehsan Amid et al.ICLR 2024 · 28 citations
Related papers
- Envisioning Class Entity Reasoning by Large Language Models for Few-shot LearningMushui Liu, Fangtai Wu, Bozheng Li, Ziqian Lu et al.AAAI 2025 · 15 citations
- VT-FSL: Bridging Vision and Text with LLMs for Few-Shot LearningWenhao Li, Qiangchang Wang, Xianjing Meng, Zhibin Wu et al.NeurIPS 2025 · 10 citations
- Multi-Label Few-Shot Image Classification via Pairwise Feature Augmentation and Flexible Prompt LearningHan Liu, Yuanyuan Wang, Xiaotong Zhang, Feng Zhang et al.AAAI 2025 · 3 citations
- Probabilistic Prototype Calibration of Vision-Language Models for Generalized Few-Shot Semantic SegmentationJie Liu, Jiayi Shen, Pan Zhou, Jan-Jakob Sonke et al.ICCV 2025 · 4 citations
- StyleProto: Style-Augmented Prototype Learning for Cross-Domain Few-Shot Object DetectionXi Yang, Quantao XieAAAI 2026
