Making Large Vision Language Models to Be Good Few-Shot Learners
Fan Liu, Wenwen Cai, Jian Huo, Chuanyi Zhang, Delong Chen, Jun Zhou
Abstract
Few-shot classification (FSC) is a fundamental yet challenging task in computer vision that involves recognizing novel classes from limited data. While previous methods have focused on enhancing visual features or incorporating additional modalities, Large Vision Language Models (LVLMs) offer a promising alternative due to their rich knowledge and strong visual perception. However, LVLMs risk learning specific response formats rather than effectively extracting useful information from support data in FSC tasks. In this paper, we investigate LVLMs' performance in FSC and identify key issues such as insufficient learning and the presence of severe positional biases. To tackle above challenges, we adopt the meta-learning strategy to teach models "learn to learn". By constructing a rich set of meta-tasks for instruction fine-tuning, LVLMs enhance the ability to extract information from few-shot support data for classification. Additionally, we further boost LVLM's few-shot learning capabilities through label augmentation and candidate selection in the fine-tuning and inference stage, respectively. Label augmentation is implemented via a character perturbation strategy to ensure the model focuses on support information. Candidate selection leverages attribute descriptions to filter out unreliable candidates and simplify the task. Extensive experiments demonstrate that our approach achieves superior performance on both general and fine-grained datasets. Furthermore, our candidate selection strategy has been proved beneficial for training-free LVLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 35955bdf-a539-4347-b54d-9db55ab123afCited by top-tier papers2
- Verbalized Representation Learning for Interpretable Few-Shot GeneralizationCheng-Fu Yang, Da Yin, Wenbo Hu, Heng Ji et al.ICCV 2025 · 1 citation
- MPA: Multimodal Prototype Augmentation for Few-Shot LearningLiwen Wu, Wei Wang, Lei Zhao, Zhan Gao et al.AAAI 2026
Builds on18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Joint Distribution Matters: Deep Brownian Distance Covariance for Few-Shot ClassificationJiangtao Xie, Fei Long, Jiaming Lv, Qilong Wang et al.CVPR 2022 · 270 citations
Related papers
- VERO: Verification and Zero-Shot Feedback Acquisition for Few-Shot Multimodal Aspect-Level Sentiment ClassificationKai Sun, Hao Wu, Bin Shi, Samuel Mensah et al.AAAI 2025 · 1 citation
- LLaFS: When Large Language Models Meet Few-Shot SegmentationLanyun Zhu, Tianrun Chen, Deyi Ji, Jieping Ye et al.CVPR 2024 · 39 citations
- Tuning Language Models as Training Data Generators for Augmentation-Enhanced Few-Shot LearningYu Meng, Martin Michalski, Jiaxin Huang, Yu Zhang et al.ICML 2023 · 64 citations
- Context-Aware Meta-LearningChristopher Fifty, Dennis Duan, Ronald G. Junkins, Ehsan Amid et al.ICLR 2024 · 28 citations
- Imagining Vision From Language for Few-Shot Class-Incremental LearningShuo Li, Xingchen Liu, Fang Liu, Licheng Jiao et al.ACM MM 2025 · 2 citations
