Better Integrating Vision and Semantics for Improving Few-shot Classification
Zhuoling Li, Yong Wang
摘要
Some recent methods address few-shot classification by integrating visual and semantic prototypes. However, they usually ignore the difference in feature structure between the visual and semantic modalities, which leads to limited performance improvements. In this paper, we propose a novel method, called bimodal integrator (BMI), to better integrate visual and semantic prototypes. In BMI, we first construct a latent space for each modality via a variational autoencoder, and then align the semantic latent space to the visual latent space. Through this semantics-to-vision alignment, the semantic modality is mapped to the visual latent space and has the same feature structure as the visual modality. As a result, the visual and semantic prototypes can be better integrated. In addition, based on the multivariate Gaussian distribution and the prompt engineering, a data augmentation scheme is designed to ensure the accuracy of modality alignment during the training process. Experimental results demonstrate that BMI significantly improves few-shot classification, making simple baselines outperform the most advanced methods on miniImageNet and tieredImageNet datasets.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- KNN Transformer with Pyramid Prompts for Few-Shot LearningWenhao Li, Qiangchang Wang, Peng Zhao, Yilong YinACM MM 2024 · 被引用 3 次
- R2-Seg: Training-Free OOD Medical Tumor Segmentation via Anatomical Reasoning and Statistical RejectionShuaike Shen, Ke Liu, Jiaqing Xie, Shangde Gao 等CVPR 2026
相关 Paper
- Generating Representative Samples for Few-Shot ClassificationJingyi Xu, Hieu LeCVPR 2022 · 被引用 96 次
- FewVS: A Vision-Semantics Integration Framework for Few-Shot Image ClassificationZhuoling Li, Yong Wang, Kaitong LiACM MM 2024 · 被引用 4 次
- Synthesized Feature based Few-Shot Class-Incremental Learning on a Mixture of SubspacesAli Cheraghian, Shafin Rahman, Sameera Ramasinghe, Pengfei Fang 等ICCV 2021 · 被引用 81 次
- VT-FSL: Bridging Vision and Text with LLMs for Few-Shot LearningWenhao Li, Qiangchang Wang, Xianjing Meng, Zhibin Wu 等NeurIPS 2025 · 被引用 10 次
- Learning to Learn Variational Semantic MemoryXiantong Zhen, Ying-Jun Du, Huan Xiong, Qiang Qiu 等NeurIPS 2020 · 被引用 40 次
