WaveDN: A Wavelet-based Training-free Zero-shot Enhancement for Vision-Language Models
Jiulin Li, Mengyu Yang, Ye Tian, Lanshan Zhang, Yongchun Lu, Jice Liu, Wendong Wang
Abstract
Vision-Language Models (VLMs) built on contrastive learning, such as CLIP, demonstrate great transferability and excel in downstream tasks like zero-shot classification and retrieval. To further enhance the performance of VLMs, existing methods have introduced additional parameter modules or fine-tuned VLMs on downstream datasets. However, these methods often fall short in scenarios where labeled data for downstream tasks is either unavailable or insufficient for fine-tuning, and the training of additional parameter modules may considerably impair the existing transferability of VLMs on open-set tasks. To alleviate this issue, we introduce WaveDN, a wavelet-based distribution normalization method that can boost the VLMs' performance on downstream tasks without parametric modules or labeled data. Initially, wavelet distributions are extracted from the embeddings of the sampled, unlabeled test samples. Subsequently, WaveDN conducts a hierarchical normalization across the wavelet coefficients of all embeddings, thereby incorporating the distributional characteristics of the test data. Finally, the normalized embeddings are reconstructed via inverse wavelet transformation, facilitating the computation of similarity metrics between the samples. Through extensive experiments on two downstream tasks, using a total of 14 datasets covering text-image and text-audio modal data, WaveDN has demonstrated superiority compared to state-of-the-art methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 06747f42-1eab-46dc-80be-c3244b97d8aeRelated papers
- Test-Time Distribution Normalization for Contrastively Learned Visual-language ModelsYifei Zhou, Juntao Ren, Fengyu Li, Ramin Zabih et al.NeurIPS 2023 · 22 citations
- Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt DiversificationYunyi Xuan, Weijie Chen, Shicai Yang, Di Xie et al.ACM MM 2023 · 4 citations
- Semi-Supervised CLIP Adaptation by Enforcing Semantic and Trapezoidal ConsistencyKai Gan, Bo Ye, Min-Ling Zhang, Tong WeiICLR 2025
- Progressive Distribution Bridging: Unsupervised Adaptation for Large-Scale Pre-Trained Models via Adaptive Auxiliary DataWeinan He, Yixin Zhang, Zilei WangICCV 2025 · 1 citation
- CLIPTTA: Robust Contrastive Vision-Language Test-Time AdaptationMarc Lafon, Gustavo Adolfo Vargas Hakim, Clément Rambour, Christian Desrosiers et al.NeurIPS 2025 · 5 citations
