DO: A Dual Debiasing Operator for Training-Free Test-Time Adaptation of Vision–Language Models
Yihong Luo, Wenwu He, Dong Liang, Yihang Zhou, Zhuo-Xu Cui
Abstract
Training-free test-time adaptation (TTA) for vision-language models (VLMs) can improve zero-shot classification under mild shifts, but often degrades under severe style/environment variation. We identify two shared failure modes in training-free pipelines: (i) retrieval confounding, where feature similarity is dominated by nuisance/style variation and corrupts retrieval evidence; and (ii) environment-biased priors, where VLM logits exhibit environment-dependent centered shifts that distort gating and prior-like terms. Therefore, we propose DO, a training-free debiasing operator that outputs three inference-time objects: a retrieval-oriented content feature for semantic matching, a style-aware routing coordinate for bias tracking, and debiased logits for corrected priors. DO composes plug-and-play with retrieval-based and closed-form Gaussian adapters in online and transductive settings. We further provide operator-to-decision guarantees: finite-difference covariance recovers a nuisance-sensitive subspace, routing-based EMA controls centered-logit bias estimates, and these errors yield bounded posterior log-odds perturbations, leading to a margin-based condition for label invariance. Extensive experiments show that DO achieves its clearest gains under style/environment-dominant shifts, with broader gains elsewhere. Code is available at https://github.com/MAiTL-Group/D2O.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f2cc356c-83b7-4cc6-bc9b-c8f70c641261Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
Related papers
- Training-Free Test-Time Adaptation via Shape and Style Guidance for Vision-Language ModelsShenglong Zhou, Manjiang Yin, Leiyu Sun, Shicai Yang et al.NeurIPS 2025 · 2 citations
- Advancing Reliable Test-Time Adaptation of Vision-Language Models under Visual VariationsYiwen Liang, Hui Chen, Yizhe Xiong, Zihan Zhou et al.ACM MM 2025 · 1 citation
- Free on the Fly: Enhancing Flexibility in Test-Time Adaptation with Online EMQiyuan Dai, Sibei YangCVPR 2025
- Backpropagation-Free Test-Time Adaptation via Probabilistic Gaussian AlignmentYoujia Zhang, Youngeun Kim, Young-Geun Choi, Hongyeob Kim et al.NeurIPS 2025 · 10 citations
- Multi-Label Test-Time Adaptation with Bayesian Conditional PriorsQiru Li, Ao Zhou, Zhiwei Jiang, Zifeng Cheng et al.ICML 2026 · 1 citation
