ViTPrompt: Training-Free Prompt Refinement with Visual Tokens for Open-Vocabulary Detection
Yitong Qin, Lihua Zhou, Jiwei Wei, Ran Ran, Shiyuan He, Zeyu Ma, Shuaifeng Li, Nianxin Li, Heng Tao Shen
Abstract
Test-Time Adaptive Object Detection (TTAOD) aims to maintain detection performance under distribution shifts without retraining. While recent vision-language models enable open-vocabulary detection, existing TTAOD methods-whether closed-set or open-vocabulary-focus solely on improving classification confidence and ignore the degradation of bounding box localization. To address this critical gap, we propose ViTPrompt (Visual Token-Prompting), a training-free framework that jointly refines both bounding boxes and class scores at test time.
Our key insight is to augment the original text prompt with instance-aware visual tokens extracted from highconfidence detections in an initial forward pass. This enhanced prompt is then fed into a second inference stage, where the cross-modal decoder produces updated box coordinates and more accurate classification logits through enriched semantic context. ViTPrompt requires no backpropagation, parameter updates, or external memory, making it highly efficient for real-time deployment. Experiments on multiple out-of-distribution benchmarks show that ViT-Prompt achieves state-of-the-art performance, with consistent gains in both localization accuracy and classification fidelity , demonstrating its effectiveness as a holistic solution for open-vocabulary TTAOD. We will release our code at https://github.com/buerzlh/Test-time-Adaptation-for-Object-Detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 74968991-e5ff-471a-aef3-f76eefcdca6eBuilds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain AdaptationJian Liang, Dapeng Hu, Jiashi FengICML 2020 · 1,624 citations
Related papers
- Test-Time Adaptive Object Detection with Foundation ModelYingjie Gao, Yanan Zhang, Zhi Cai, Di HuangNeurIPS 2025 · 7 citations
- Scene-adaptive and Region-aware Multi-modal Prompt for Open Vocabulary Object DetectionXiaowei Zhao, Xianglong Liu, Duorui Wang, Yajun Gao et al.CVPR 2024 · 8 citations
- Flatness Guided Test-Time Adaptation for Vision-Language ModelsAodi Li, Liansheng Zhuang, Xiao Long, Houqiang Li et al.ICLR 2026 · 1 citation
- Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language ModelYu Du, Fangyun Wei, Zihe Zhang, Miaojing Shi et al.CVPR 2022 · 311 citations
- VPA: Fully Test-Time Visual Prompt AdaptationJiachen Sun, Mark Ibrahim, Melissa Hall, Ivan Evtimov et al.ACM MM 2023 · 7 citations
