Lune

AAAI2025顶会

Black-Box Test-Time Prompt Tuning for Vision-Language Models

Fan'an Meng, Chaoran Cui, Hongjun Dai, Shuai Gong

2025年份
6被引次数
6顶会引用

摘要

Test-time prompt tuning (TPT) aims to adjust the visionlanguage models (e.g., CLIP) with learnable prompts during the inference phase. However, previous works overlooked that pre-trained models as a service (MaaS) have become a noticeable trend due to their commercial usage and potential risk of misuse. In the context of MaaS, users can only design prompts in inputs and query the black-box visionlanguage models through inference APIs, rendering the previous paradigm of utilizing gradient for prompt tuning is infeasible. In this paper, we propose black-box test-time prompt tuning (B 2 TPT), a novel framework that addresses the challenge of optimizing prompts without gradients in an unsupervised manner. Specifically, B 2 TPT designs a consistent or confident (CoC) pseudo-labeling strategy to generate highquality pseudo-labels from the outputs. Subsequently, we propose to optimize low-dimensional intrinsic prompts using a derivative-free evolution algorithm and to project them onto the original text and vision prompts. This strategy addresses the gradient-free challenge while reducing complexity. Extensive experiments across 15 datasets demonstrate the superiority of B 2 TPT. The results show that B 2 TPT not only outperforms CLIP's zero-shot inference at test time, but also surpasses other gradient-based TPT methods.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper6

问问它们各自怎么用它

它引用的顶会 Paper18

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖