ViTPrompt: Training-Free Prompt Refinement with Visual Tokens for Open-Vocabulary Detection
Yitong Qin, Lihua Zhou, Jiwei Wei, Ran Ran, Shiyuan He, Zeyu Ma, Shuaifeng Li, Nianxin Li, Heng Tao Shen
摘要
Test-Time Adaptive Object Detection (TTAOD) aims to maintain detection performance under distribution shifts without retraining. While recent vision-language models enable open-vocabulary detection, existing TTAOD methods-whether closed-set or open-vocabulary-focus solely on improving classification confidence and ignore the degradation of bounding box localization. To address this critical gap, we propose ViTPrompt (Visual Token-Prompting), a training-free framework that jointly refines both bounding boxes and class scores at test time.
Our key insight is to augment the original text prompt with instance-aware visual tokens extracted from highconfidence detections in an initial forward pass. This enhanced prompt is then fed into a second inference stage, where the cross-modal decoder produces updated box coordinates and more accurate classification logits through enriched semantic context. ViTPrompt requires no backpropagation, parameter updates, or external memory, making it highly efficient for real-time deployment. Experiments on multiple out-of-distribution benchmarks show that ViT-Prompt achieves state-of-the-art performance, with consistent gains in both localization accuracy and classification fidelity , demonstrating its effectiveness as a holistic solution for open-vocabulary TTAOD. We will release our code at https://github.com/buerzlh/Test-time-Adaptation-for-Object-Detection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen 等ICLR 2021 · 被引用 1,731 次
- Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain AdaptationJian Liang, Dapeng Hu, Jiashi FengICML 2020 · 被引用 1,624 次
相关 Paper
- Test-Time Adaptive Object Detection with Foundation ModelYingjie Gao, Yanan Zhang, Zhi Cai, Di HuangNeurIPS 2025 · 被引用 7 次
- Scene-adaptive and Region-aware Multi-modal Prompt for Open Vocabulary Object DetectionXiaowei Zhao, Xianglong Liu, Duorui Wang, Yajun Gao 等CVPR 2024 · 被引用 8 次
- Flatness Guided Test-Time Adaptation for Vision-Language ModelsAodi Li, Liansheng Zhuang, Xiao Long, Houqiang Li 等ICLR 2026 · 被引用 1 次
- Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language ModelYu Du, Fangyun Wei, Zihe Zhang, Miaojing Shi 等CVPR 2022 · 被引用 311 次
- VPA: Fully Test-Time Visual Prompt AdaptationJiachen Sun, Mark Ibrahim, Melissa Hall, Ivan Evtimov 等ACM MM 2023 · 被引用 7 次
