Test-Time Adaptation of Vision-Language Models for Open-Vocabulary Semantic Segmentation
Mehrdad Noori, David Osowiechi, Gustavo Adolfo Vargas Hakim, Ali Bahri, Moslem Yazdanpanah, Sahar Dastani, Farzad Beizaee, Ismail Ben Ayed, Christian Desrosiers
Abstract
Recently, test-time adaptation has attracted wide interest in the context of vision-language models for image classification. However, to the best of our knowledge, the problem is completely overlooked in dense prediction tasks such as Open-Vocabulary Semantic Segmentation (OVSS). In response, we propose a novel TTA method tailored to adapting VLMs for segmentation during test time. Unlike TTA methods for image classification, our Multi-Level and Multi-Prompt (MLMP) entropy minimization integrates features from intermediate vision-encoder layers and is performed with different text-prompt templates at both the global CLS token and local pixel-wise levels. Our approach could be used as plug-and-play for any segmentation network, does not require additional training data or labels, and remains effective even with a single test sample. Furthermore, we introduce a comprehensive OVSS TTA benchmark suite, which integrates a rigorous evaluation protocol, nine segmentation datasets, 15 common synthetic corruptions, and additional real and rendered domain shifts, with a total of 87 distinct test scenarios, establishing a standardized and comprehensive testbed for future TTA research in open-vocabulary segmentation. Our experiments on this suite demonstrate that our segmentation-tailored method consistently delivers significant gains over direct adoption of TTA classification baselines. Code and data are available at https://github.com/dosowiechi/MLMP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf3ce422-6711-4856-9c00-ed0a5d11464fCited by top-tier papers5
- Purge-Gate: Backpropagation-Free Test-Time Adaptation for Point Clouds Classification via Token PurgingMoslem Yazdanpanah, Ali Bahri, Mehrdad Noori, Sahar Dastani et al.ICCV 2025 · 3 citations
- StructXLIP: Enhancing Vision-language Models with Multimodal Structural CuesZanxi Ruan, Songqun Gao, Qiuyu Kong, Yiming Wang et al.CVPR 2026 · 1 citation
- Test-Time Multi-Prompt Adaptation for Open-Vocabulary Remote Sensing Image SegmentationTing Yang, Qilong Wang, Qibin Hou, Qinghua HuCVPR 2026
- Retrieve and Segment: Are a Few Examples Enough to Bridge the Supervision Gap in Open-Vocabulary Segmentation?Tilemachos Aravanis, Vladan Stojnic, Bill Psomas, Nikos Komodakis et al.CVPR 2026
- TRUST: Test-Time Refinement using Uncertainty-Guided SSM TraversesSahar Dastani, Ali Bahri, Gustavo Adolfo Vargas Hakim, Moslem Yazdanpanah et al.NeurIPS 2025
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- PANet: Few-Shot Image Semantic Segmentation With Prototype AlignmentKaixin Wang, Jun Hao Liew, Yingtian Zou, Daquan Zhou et al.ICCV 2019 · 1,404 citations
- ACDC: The Adverse Conditions Dataset with Correspondences for Semantic Driving Scene UnderstandingChristos Sakaridis, Dengxin Dai, Luc Van GoolICCV 2021 · 655 citations
- Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language ModelsManli Shu, Weili Nie, De-An Huang, Zhiding Yu et al.NeurIPS 2022 · 603 citations
Related papers
- Dual Prototype Evolving for Test-Time Generalization of Vision-Language ModelsCe Zhang, Simon Stepputtis, Katia P. Sycara, Yaqi XieNeurIPS 2024 · 57 citations
- PatAug: Augmentation of Augmentation for Test-Time AdaptationXinyao Li, Dan Zhang, Zhekai Du, Lei Zhu et al.ACM MM 2025 · 2 citations
- Free on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EMQiyuan Dai, Sibei YangCVPR 2025
- CLIPTTA: Robust Contrastive Vision-Language Test-Time AdaptationMarc Lafon, Gustavo Adolfo Vargas Hakim, Clément Rambour, Christian Desrosiers et al.NeurIPS 2025 · 5 citations
- WATT: Weight Average Test Time Adaptation of CLIPDavid Osowiechi, Mehrdad Noori, Gustavo Adolfo Vargas Hakim, Moslem Yazdanpanah et al.NeurIPS 2024 · 46 citations
