CASPA: Graph-Structured Concept Anchors for Modality-Agnostic Adaptation in Vision-Language Models
Abhiroop Chatterjee, Susmita Ghosh, Ashish Ghosh, Emmett J. Ientilucci
摘要
Recent advances in vision-language models (VLMs) have revealed both the promise and the rigidity of large-scale pretraining. Despite their impressive zero-shot generalization, existing adaptation paradigms-whether prompttuning, adapter injection, or fine-tuning-remain classspecific, modality-biased, and structure-agnostic. However, these design choices limit reasoning-level transfer across tasks. To this end, we rethink adaptation as a shared conceptual structure rather than a per-class specialization. We propose CASPA (Concept-Anchored Semantic Prompt Adapter), a dual-anchor semantic adapter that jointly learns shared text and image anchors as a bidirectional conceptual interface between modalities. Each class learns a soft association distribution over these anchors, producing compositional representations which enable parameter sharing and semantic reuse. To further align visual and textual reasoning spaces, CASPA employs Semantic Cross-Consistency Regularization (S-XCR), enforcing geometric and semantic agreement between text-and imageconditioned anchor mixtures. CASPA, therefore, provides a structurally constrained alternative to class-conditional prompt parameterization while keeping the CLIP backbone frozen. Evaluated across four regimes, Base-to-Novel setup, cross-data transfer, few-shot, and backbone-agnostic evaluations, on eleven diverse visual recognition datasets, CASPA matches or outperforms state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan 等NeurIPS 2020 · 被引用 1,631 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
相关 Paper
- APoLLo : Unified Adapter and Prompt Learning for Vision Language ModelsSanjoy Chowdhury, Sayan Nag, Dinesh ManochaEMNLP 2023 · 被引用 17 次
- Semi-Supervised CLIP Adaptation by Enforcing Semantic and Trapezoidal ConsistencyKai Gan, Bo Ye, Min-Ling Zhang, Tong WeiICLR 2025
- MaPLe: Multi-modal Prompt LearningMuhammad Uzair Khattak, Hanoona Abdul Rasheed, Muhammad Maaz, Salman H. Khan 等CVPR 2023
- Advancing Prompt Learning through an External LayerFangming Cui, Xun Yang, Chao Wu, Liang Xiao 等ACM MM 2024 · 被引用 3 次
- ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided OptimizationHao Wang, Fang Liu, Licheng Jiao, Jiahao Wang 等AAAI 2024 · 被引用 54 次
