Homology Consistency Constrained Efficient Tuning for Vision-Language Models
Huatian Zhang, Lei Zhang, Yongdong Zhang, Zhendong Mao
Abstract
Efficient transfer learning has shown remarkable performance in tuning large-scale vision-language models (VLMs) toward downstream tasks with limited data resources. The key challenge of efficient transfer lies in adjusting image-text alignment to be task-specific while preserving pre-trained general knowledge. However, existing methods adjust image-text alignment merely on a set of observed samples, e.g. , data set and external knowledge base, which cannot guarantee to keep the correspondence of general concepts between image and text latent manifolds without being disrupted and thereby a weak generalization of the adjusted alignment. In this work, we propose a Homology Consistency (HC) constraint for efficient transfer on VLMs, which explicitly constrains the correspondence of image and text latent manifolds through structural equivalence based on persistent homology in downstream tuning. Specifically, we build simplicial complex on the top of data to mimic the topology of latent manifolds, then track the persistence of the homology classes of topological features across multiple scales, and guide the directions of persistence tracks in image and text manifolds to coincide each other, with a deviating perturbation additionally. For practical application, we tailor the implementation of our proposed HC constraint for two main paradigms of adapter tuning. Extensive experiments on few-shot learning over 11 datasets and domain generalization demonstrate the effectiveness and robustness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b691f18b-0175-41dd-a9e0-0129e49f1455Cited by top-tier papers1
Ask how each one uses itBuilds on37
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- Preserve and Sculpt: Manifold-Aligned Fine-tuning of Vision-Language Models for Few-Shot LearningDexia Chen, Qianjie Zhu, Weibing Li, Yue Yu et al.ICLR 2026
- Prompt Tuning for CLIP on the Pretrained ManifoldXi Yang, Yuanrong Xu, Weigang Zhang, Guangming Lu et al.ICML 2026 · 1 citation
- Task Residual for Tuning Vision-Language ModelsTao Yu, Zhihe Lu, Xin Jin, Zhibo Chen et al.CVPR 2023
- Efficient Transfer Learning for Video-language Foundation ModelsHaoxing Chen, Zizheng Huang, Yan Hong, Yanshuo Wang et al.CVPR 2025
- MMA: Multi-Modal Adapter for Vision-Language ModelsLingxiao Yang, Ru-Yuan Zhang, Yanchen Wang, Xiaohua XieCVPR 2024 · 46 citations
