ZIP: An Efficient Zeroth-order Prompt Tuning for Black-box Vision-Language Models
Seonghwan Park, Jaehyeon Jeong, Yongjun Kim, Jaeho Lee, Namhoon Lee
Abstract
Recent studies have introduced various approaches for prompt-tuning black-box vision-language models, referred to as black-box prompt-tuning (BBPT). While BBPT has demonstrated considerable potential, it is often found that many existing methods require an excessive number of queries (i.e., function evaluations), which poses a significant challenge in real-world scenarios where the number of allowed queries is limited. To tackle this issue, we propose Zeroth-order Intrinsic-dimensional Prompt-tuning (ZIP), a novel approach that enables efficient and robust prompt optimization in a purely black-box setting. The key idea of ZIP is to reduce the problem dimensionality and the variance of zeroth-order gradient estimates, such that the training is done fast with far less queries. We achieve this by re-parameterizing prompts in low-rank representations and designing intrinsic-dimensional clipping of estimated gradients. We evaluate ZIP on 13+ vision-language tasks in standard benchmarks and show that it achieves an average improvement of approximately 6% in few-shot accuracy and 48% in query efficiency compared to the best-performing alternative BBPT methods, establishing a new state of the art. Our ablation analysis further shows that the proposed clipping mechanism is robust and nearly optimal, without the need to manually select the clipping threshold, matching the result of expensive hyperparameter search.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f5ce162-9422-4522-a067-2d949a3a19f6Cited by top-tier papers6
- Prime Once, then Reprogram Locally: An Efficient Alternative to Black-Box Service Model AdaptationYunbei Zhang, Chengyi Cai, Feng Liu, Jihun HammCVPR 2026 · 5 citations
- SharpZO: Hybrid Sharpness-Aware Vision Language Model Prompt Tuning via Forward-Only PassesYifan Yang, Zhen Zhang, Rupak Vignesh Swaminathan, Jing Liu et al.NeurIPS 2025 · 4 citations
- Federated Learning with Unlabeled Clients: Personalization Can Happen in Low DimensionsHossein Zakerinia, Jonathan Scott, Christoph LampertICML 2026
- ZOO-Prune: Training-Free Token Pruning via Zeroth-Order Gradient Estimation in Vision-Language ModelsYoungeun Kim, Youjia Zhang, Huiling Liu, Aecheon Jung et al.CVPR 2026
- VALIANT: Prompt Instability for Active Learning in Black-Box Medical ImagingDwarikanath Mahapatra, Behzad Bozorgtabar, Sudipta Roy, Imran Razzak et al.AAAI 2026
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- PromptBoosting: Black-Box Text Classification with Ten Forward PassesBairu Hou, Joe O'Connor, Jacob Andreas, Shiyu Chang et al.ICML 2023 · 56 citations
- Query Efficient Black-Box Visual Prompting with Subspace LearningZhaogeng Liu, Haozhen Zhang, Hualin Zhang, Xingchen Li et al.CVPR 2025
- Subspace Selection based Prompt Tuning with Nonconvex Nonsmooth Black-Box OptimizationHaozhen Zhang, Hualin Zhang, Bin Gu, Yi ChangKDD 2024 · 1 citation
- Black-Box Test-Time Prompt Tuning for Vision-Language ModelsFan'an Meng, Chaoran Cui, Hongjun Dai, Shuai GongAAAI 2025 · 6 citations
- Modality-Agnostic Zeroth-Order LoRA Fine-Tuning for Black-Box Prompt OptimizationXingchen Li, Jia Zhang, Tianxing Man, Wenkang Wang et al.KDD 2026
