VT-FSL: Bridging Vision and Text with LLMs for Few-Shot Learning
Wenhao Li, Qiangchang Wang, Xianjing Meng, Zhibin Wu, Yilong Yin
摘要
Few-shot learning (FSL) aims to recognize novel concepts from only a few labeled support samples. Recent studies enhance support features by incorporating additional semantic information (e.g., class descriptions) or designing complex semantic fusion modules. However, these methods still suffer from hallucinating semantics that contradict the visual evidence due to the lack of grounding in actual instances, resulting in noisy guidance and costly corrections. To address these issues, we propose a novel framework, bridging Vision and Text with LLMs for Few-Shot Learning (VT-FSL), which constructs precise cross-modal prompts conditioned on Large Language Models (LLMs) and support images, seamlessly integrating them through a geometry-aware alignment mechanism. It mainly consists of Cross-modal Iterative Prompting (CIP) and Cross-modal Geometric Alignment (CGA). Specifically, the CIP conditions an LLM on both class names and support images to generate precise class descriptions iteratively in a single structured reasoning pass. These descriptions not only enrich the semantic understanding of novel classes but also enable the zero-shot synthesis of semantically consistent images. The descriptions and synthetic images act respectively as complementary textual and visual prompts, providing high-level class semantics and low-level intra-class diversity to compensate for limited support data. Furthermore, the CGA jointly aligns the fused textual, support, and synthetic visual representations by minimizing the kernelized volume of the 3-dimensional parallelotope they span. It captures global and nonlinear relationships among all representations, enabling structured and consistent multimodal integration. The proposed VT-FSL method establishes new state-of-the-art performance across ten diverse benchmarks, including standard, cross-domain, and fine-grained few-shot learning scenarios. Code is available at https://github.com/peacelwh/VT-FSL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- DVLA-RL: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot LearningWenhao Li, Xianjing Meng, Qiangchang Wang, Zhongyi Han 等ICLR 2026 · 被引用 4 次
- From Pretraining to Pathology: How Noise Leads to Catastrophic Inheritance in Medical ModelsHao Sun, Zhongyi Han, Hao Chen, Jindong Wang 等NeurIPS 2025 · 被引用 4 次
- CAST-LUT: Tokenizer-Guided HSV Look-Up Tables for Purple Flare RemovalPu Wang, Shuning Sun, Jialang Lu, Chen Wu 等AAAI 2026 · 被引用 2 次
- Cross-View Lewis Weight Fusion Empowering Exemplar Replay for Federated Class-Incremental LearningZhuang Qi, Yingpeng Tang, Lei Meng, Xiaoxiao Li 等ICML 2026
- Multimodal Causality-Driven Representation Learning for Generalizable Medical Image SegmentationXUSHENG LIANG, Lihua Zhou, Nianxin Li, miao xu 等CVPR 2026
它引用的顶会 Paper48
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Prompt-aligned Gradient for Prompt TuningBeier Zhu, Yulei Niu, Yucheng Han, Yue Wu 等ICCV 2023 · 被引用 475 次
相关 Paper
- DSV-LFS: Unifying LLM-Driven Semantic Cues with Visual Features for Robust Few-Shot SegmentationAmin Karimi, Charalambos PoullisCVPR 2025
- Text Augmented Correlation Transformer For Few-shot Classification & SegmentationSrinivasa Rao Nandam, Sara Atito, Zhenhua Feng, Josef Kittler 等CVPR 2025
- Envisioning Class Entity Reasoning by Large Language Models for Few-shot LearningMushui Liu, Fangtai Wu, Bozheng Li, Ziqian Lu 等AAAI 2025 · 被引用 15 次
- Semantic-Guided Global-Local Collaborative Prompt Learning for Few-Shot Class Incremental Learningyongxin yan, Weisen Chen, Xingye Chen, Yuanjie Shao 等CVPR 2026
- Large Language Models are Good Prompt Learners for Low-Shot Image ClassificationZhaoheng Zheng, Jingmin Wei, Xuefeng Hu, Haidong Zhu 等CVPR 2024 · 被引用 15 次
