Lune

AAAI2025Top-tier venue

Black-Box Test-Time Prompt Tuning for Vision-Language Models

Fan'an Meng, Chaoran Cui, Hongjun Dai, Shuai Gong

2025Year
6Citations
6Top-tier citations

Abstract

Test-time prompt tuning (TPT) aims to adjust the visionlanguage models (e.g., CLIP) with learnable prompts during the inference phase. However, previous works overlooked that pre-trained models as a service (MaaS) have become a noticeable trend due to their commercial usage and potential risk of misuse. In the context of MaaS, users can only design prompts in inputs and query the black-box visionlanguage models through inference APIs, rendering the previous paradigm of utilizing gradient for prompt tuning is infeasible. In this paper, we propose black-box test-time prompt tuning (B 2 TPT), a novel framework that addresses the challenge of optimizing prompts without gradients in an unsupervised manner. Specifically, B 2 TPT designs a consistent or confident (CoC) pseudo-labeling strategy to generate highquality pseudo-labels from the outputs. Subsequently, we propose to optimize low-dimensional intrinsic prompts using a derivative-free evolution algorithm and to project them onto the original text and vision prompts. This strategy addresses the gradient-free challenge while reducing complexity. Extensive experiments across 15 datasets demonstrate the superiority of B 2 TPT. The results show that B 2 TPT not only outperforms CLIP's zero-shot inference at test time, but also surpasses other gradient-based TPT methods.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 841919ed-85d4-4907-a395-cf3f8c2cea30

Cited by top-tier papers6

Ask how each one uses it

Builds on18

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines