Better Integrating Vision and Semantics for Improving Few-shot Classification
Zhuoling Li, Yong Wang
Abstract
Some recent methods address few-shot classification by integrating visual and semantic prototypes. However, they usually ignore the difference in feature structure between the visual and semantic modalities, which leads to limited performance improvements. In this paper, we propose a novel method, called bimodal integrator (BMI), to better integrate visual and semantic prototypes. In BMI, we first construct a latent space for each modality via a variational autoencoder, and then align the semantic latent space to the visual latent space. Through this semantics-to-vision alignment, the semantic modality is mapped to the visual latent space and has the same feature structure as the visual modality. As a result, the visual and semantic prototypes can be better integrated. In addition, based on the multivariate Gaussian distribution and the prompt engineering, a data augmentation scheme is designed to ensure the accuracy of modality alignment during the training process. Experimental results demonstrate that BMI significantly improves few-shot classification, making simple baselines outperform the most advanced methods on miniImageNet and tieredImageNet datasets.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3a0f5032-6828-402f-b3f3-2b45b7754e50Cited by top-tier papers2
- KNN Transformer with Pyramid Prompts for Few-Shot LearningWenhao Li, Qiangchang Wang, Peng Zhao, Yilong YinACM MM 2024 · 3 citations
- R2-Seg: Training-Free OOD Medical Tumor Segmentation via Anatomical Reasoning and Statistical RejectionShuaike Shen, Ke Liu, Jiaqing Xie, Shangde Gao et al.CVPR 2026
Related papers
- Generating Representative Samples for Few-Shot ClassificationJingyi Xu, Hieu LeCVPR 2022 · 96 citations
- FewVS: A Vision-Semantics Integration Framework for Few-Shot Image ClassificationZhuoling Li, Yong Wang, Kaitong LiACM MM 2024 · 4 citations
- Synthesized Feature based Few-Shot Class-Incremental Learning on a Mixture of SubspacesAli Cheraghian, Shafin Rahman, Sameera Ramasinghe, Pengfei Fang et al.ICCV 2021 · 81 citations
- VT-FSL: Bridging Vision and Text with LLMs for Few-Shot LearningWenhao Li, Qiangchang Wang, Xianjing Meng, Zhibin Wu et al.NeurIPS 2025 · 10 citations
- Learning to Learn Variational Semantic MemoryXiantong Zhen, Ying-Jun Du, Huan Xiong, Qiang Qiu et al.NeurIPS 2020 · 40 citations
