CIA: Class- and Instance-aware Adaptation for Vision-Language Models
Lin Peng, Cong Wan, Shaokun Wang, Xiang Song, Yuhang He, Yihong Gong
Abstract
Few-shot parameter-efficient tuning methods demonstrate promising potential for Vision-Language (V-L) models in downstream tasks. However, existing approaches primarily focus on class-level alignment between image and text features, overlooking crucial instance-specific semantic information. This limitation leads to suboptimal performance on challenging tasks and restricted generalization capability to unseen data. To address these issues, we propose Class- and Instance-aware Adaptation (CIA), a novel framework that simultaneously optimizes both class-level and instance-level alignments. Specifically, CIA introduces a novel instance encoder that leverages cross-modal self-attention to generate instance-specific text features, accompanied by a carefully designed regularization mechanism to maintain consistency between class-level and instance-level representations. Extensive experiments across 15 benchmark datasets demonstrate that CIA significantly improves the downstream adaptation of V-L models.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- Shared & Domain Self-Adaptive Experts with Frequency-Aware Discrimination for Continual Test-Time AdaptationJianchao Zhao, Chenhao Ding, Songlin Dong, Jiangyang Li et al.AAAI 2026 · 1 citation
- StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video RetrievalShaokun Wang, Weili Guan, Jizhou Han, Jianlong Wu et al.SIGIR 2026
- Learning Like Humans: Analogical Concept Learning for Generalized Category DiscoveryJizhou Han, Chenhao Ding, Yuhang He, Qiang Wang et al.CVPR 2026
Related papers
- CASPA: Graph-Structured Concept Anchors for Modality-Agnostic Adaptation in Vision-Language ModelsAbhiroop Chatterjee, Susmita Ghosh, Ashish Ghosh, Emmett J. IentilucciCVPR 2026
- Fine-Grained Visual Prompt Learning of Vision-Language Models for Image RecognitionHongbo Sun, Xiangteng He, Jiahuan Zhou, Yuxin PengACM MM 2023 · 16 citations
- LaViP: Language-Grounded Visual PromptingNilakshan Kunananthaseelan, Jing Zhang, Mehrtash HarandiAAAI 2024 · 6 citations
- APoLLo : Unified Adapter and Prompt Learning for Vision Language ModelsSanjoy Chowdhury, Sayan Nag, Dinesh ManochaEMNLP 2023 · 17 citations
- InsAT: Instance-aware Semantic Alignment and Transfer from Human-Object Keypoints for Zero-to-Few-shot Action UnderstandingKazuki TsutsukawaACL 2026
