DOTA: Distributional Test-time Adaptation of Vision-Language Models
Zongbo Han, Jialong Yang, Guangyu Wang, Junfan Li, Qianli Xu, Mike Zheng Shou, Changqing Zhang
Abstract
Vision-language foundation models (VLMs), such as CLIP, exhibit remarkable performance across a wide range of tasks. However, deploying these models can be unreliable when significant distribution gaps exist between training and test data, while fine-tuning for diverse scenarios is often costly. This creates a need for methods that can efficiently adapt to new data at test time without expensive retraining. Cache-based test-time adapters serve this purpose by storing representative test samples to guide subsequent classifications. Yet, these methods typically employ naive cache management with limited capacity, leading to severe catastrophic forgetting when samples are inevitably dropped during updates. In this paper, we propose Dota (DistributiOnal Test-time Adaptation), a simple yet effective method addressing this limitation. Crucially, instead of merely memorizing individual test samples, Dota continuously estimates the underlying distribution of the test data stream. Test-time posterior probabilities are then computed using these dynamically estimated distributions via Bayes' theorem for adaptation. This distribution-centric approach enables the model to continually learn and adapt to the deployment environment. Extensive experiments validate that Dota significantly mitigates forgetting and achieves state-of-the-art performance compared to existing methods. Code is available at https://github.com/skylineeeeen/DOTA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da28cf42-5c8b-49bd-84fd-1ee6ac853b19Cited by top-tier papers10
- Backpropagation-Free Test-Time Adaptation via Probabilistic Gaussian AlignmentYoujia Zhang, Youngeun Kim, Young-Geun Choi, Hongyeob Kim et al.NeurIPS 2025 · 10 citations
- Statistics Caching Test-Time Adaptation for Vision-Language ModelsZenghao Guan, Yucan Zhou, Wu Liu, Xiaoyan GuNeurIPS 2025 · 5 citations
- Multi-Label Test-Time Adaptation with Bayesian Conditional PriorsQiru Li, Ao Zhou, Zhiwei Jiang, Zifeng Cheng et al.ICML 2026 · 1 citation
- Target-Agnostic Calibration under Distribution Shift with Frequency-Aware Gradient RectificationYilin Zhang, Cai Xu, You Wu, Ziyu Guan et al.ICML 2026 · 1 citation
- Adapting Point Cloud Analysis via Multimodal Bayesian Distribution LearningXingyu Zhu, Yi Liang, Shuo Wang, Wenbo Zhu et al.CVPR 2026
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
Related papers
- PAF: Prototype Adaptive Fusion for Test-Time Adaptation of Vision-Language ModelsSi Chen, Yujia Chen, Xiaotian Yin, Xin Liu et al.ACM MM 2025 · 1 citation
- Test-Time Retrieval-Augmented Adaptation for Vision-Language ModelsXinqi Fan, Xueli Chen, Luoxiao Yang, Chuin Hong Yap et al.ICCV 2025 · 4 citations
- DART: Dual-Modal Adaptive Online Prompting and Knowledge Retention for Test-Time AdaptationZichen Liu, Hongbo Sun, Yuxin Peng, Jiahuan ZhouAAAI 2024 · 14 citations
- Prototype-Based Test-Time Adaptation of Vision-Language ModelsZhaohong Huang, Yuxin Zhang, Wenjing Liu, Fei Chao et al.ICML 2026
- Advancing Reliable Test-Time Adaptation of Vision-Language Models under Visual VariationsYiwen Liang, Hui Chen, Yizhe Xiong, Zihan Zhou et al.ACM MM 2025 · 1 citation
