DVLA-RL: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot Learning
Wenhao Li, Xianjing Meng, Qiangchang Wang, Zhongyi Han, Zhibin Wu, Yilong Yin
Abstract
Few-shot learning (FSL) aims to generalize to novel categories with only a few samples. Recent approaches incorporate large language models (LLMs) to enrich visual representations with semantic embeddings derived from class names. However, they overlook progressive and adaptive alignment between vision and language from low-level to high-level semantics, resulting in limited semantic gains. To address these challenges, we propose Dual-level Vision–Language Alignment with Reinforcement Learning gating (DVLA-RL), which consists of Dual-level Semantic Construction (DSC) and RL-gated Attention (RLA). Specifically, DSC conditions LLMs on both class names and support samples to generate discriminative attributes, progressively selects the most relevant ones, and then synthesizes them into coherent class descriptions. This process provides complementary low-level attributes and high-level descriptions, enabling both fine-grained grounding and holistic class understanding. To dynamically integrate dual-level semantics along with the visual network layers, RLA formulates cross-modal fusion as a sequential decision process. A lightweight policy trained with episodic REINFORCE adaptively adjusts the contributions of self-attention and cross-attention to integrate textual and visual tokens. As a result, shallow layers refine local attributes and deep layers emphasize global semantics, enabling more precise cross-modal alignment. This achieves class-specific discrimination and generalized representations with merely a few support samples. DVLA-RL achieves new state-of-the-art performance across nine benchmarks in three diverse FSL scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d4ece8a-de8b-4b05-8ae7-a7b5aec65f4bCited by top-tier papers3
- Cross-View Lewis Weight Fusion Empowering Exemplar Replay for Federated Class-Incremental LearningZhuang Qi, Yingpeng Tang, Lei Meng, Xiaoxiao Li et al.ICML 2026
- Multimodal Causality-Driven Representation Learning for Generalizable Medical Image SegmentationXUSHENG LIANG, Lihua Zhou, Nianxin Li, miao xu et al.CVPR 2026
- From Selection to Scheduling: Federated Geometry-Aware Correction Makes Exemplar Replay Work Better under Continual Dynamic HeterogeneityZhuang Qi, Ying-Peng Tang, Lei Meng, Guoqing Chao et al.CVPR 2026
Builds on36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Cross-Domain Few-Shot Classification via Learned Feature-Wise TransformationHung-Yu Tseng, Hsin-Ying Lee, Jia-Bin Huang, Ming-Hsuan YangICLR 2020 · 467 citations
- Visformer: The Vision-friendly TransformerZhengsu Chen, Lingxi Xie, Jianwei Niu, Xuefeng Liu et al.ICCV 2021 · 293 citations
- Rethinking Generalization in Few-Shot ClassificationMarkus Hiller, Rongkai Ma, Mehrtash Harandi, Tom DrummondNeurIPS 2022 · 119 citations
Related papers
- VT-FSL: Bridging Vision and Text with LLMs for Few-Shot LearningWenhao Li, Qiangchang Wang, Xianjing Meng, Zhibin Wu et al.NeurIPS 2025 · 10 citations
- DSV-LFS: Unifying LLM-Driven Semantic Cues with Visual Features for Robust Few-Shot SegmentationAmin Karimi, Charalambos PoullisCVPR 2025
- Envisioning Class Entity Reasoning by Large Language Models for Few-shot LearningMushui Liu, Fangtai Wu, Bozheng Li, Ziqian Lu et al.AAAI 2025 · 15 citations
- Semantic-Guided Global-Local Collaborative Prompt Learning for Few-Shot Class Incremental Learningyongxin yan, Weisen Chen, Xingye Chen, Yuanjie Shao et al.CVPR 2026
- Language Does Matter for Cross-Domain Few-Shot Visual Feature EnhancementFei Zhou, Xiwen Zhang, Qingqing Qiu, Lei Zhang et al.CVPR 2026
