Lune

CVPR2026顶会

ViTPrompt: Training-Free Prompt Refinement with Visual Tokens for Open-Vocabulary Detection

Yitong Qin, Lihua Zhou, Jiwei Wei, Ran Ran, Shiyuan He, Zeyu Ma, Shuaifeng Li, Nianxin Li, Heng Tao Shen

出版方
2026年份

摘要

Test-Time Adaptive Object Detection (TTAOD) aims to maintain detection performance under distribution shifts without retraining. While recent vision-language models enable open-vocabulary detection, existing TTAOD methods-whether closed-set or open-vocabulary-focus solely on improving classification confidence and ignore the degradation of bounding box localization. To address this critical gap, we propose ViTPrompt (Visual Token-Prompting), a training-free framework that jointly refines both bounding boxes and class scores at test time.

Our key insight is to augment the original text prompt with instance-aware visual tokens extracted from highconfidence detections in an initial forward pass. This enhanced prompt is then fed into a second inference stage, where the cross-modal decoder produces updated box coordinates and more accurate classification logits through enriched semantic context. ViTPrompt requires no backpropagation, parameter updates, or external memory, making it highly efficient for real-time deployment. Experiments on multiple out-of-distribution benchmarks show that ViT-Prompt achieves state-of-the-art performance, with consistent gains in both localization accuracy and classification fidelity , demonstrating its effectiveness as a holistic solution for open-vocabulary TTAOD. We will release our code at https://github.com/buerzlh/Test-time-Adaptation-for-Object-Detection.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 74968991-e5ff-471a-aef3-f76eefcdca6e

它引用的顶会 Paper32

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖