Language Models as Black-Box Optimizers for Vision-Language Models
Shihong Liu, Samuel Yu, Zhiqiu Lin, Deepak Pathak, Deva Ramanan
Abstract
Vision-language models (VLMs) pre-trained on web-scale datasets have demonstrated remarkable capabilities on downstream tasks when fine-tuned with minimal data. However, many VLMs rely on proprietary data and are not open-source, which restricts the use of white-box approaches for fine-tuning. As such, we aim to develop a black-box approach to optimize VLMs through natural language prompts, thereby avoiding the need to access model parameters, feature embeddings, or even output logits. We propose employing chat-based LLMs to search for the best text prompt for VLMs. Specifically, we adopt an automatic “hill-climbing” procedure that converges to an effective prompt by evaluating the performance of current prompts and asking LLMs to refine them based on textual feedback, all within a conversational process without human-in-the-loop. In a challenging 1-shot image classification setup, our simple approach surpasses the white-box continuous prompting method (CoOp) by an average of1.5% across 11 datasets including ImageNet. Our approach also outperforms both human-engineered and LLM-generated prompts. We high-light the advantage of conversational feedback that incor-porates both positive and negative prompts, suggesting that LLMs can utilize the implicit “gradient” direction in textual feedback for a more efficient search. In addition, we find that the text prompts generated through our strategy are not only more interpretable but also transfer well across different VLM architectures in a black-box manner. Lastly, we demonstrate our framework on a state-of-the-art black-box VLM (DALL-E 3) for text-to-image optimization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c66ed707-7382-4c25-99d2-3a07b059a01cCited by top-tier papers15
- Efficient and Accurate Prompt Optimization: the Benefit of Memory in Exemplar-Guided ReflectionCilin Yan, Jingyun Wang, Lin Zhang, Ruihui Zhao et al.ACL 2025 · 16 citations
- The Neglected Tails in Vision-Language ModelsShubham Parashar, Zhiqiu Lin, Tian Liu, Xiangjue Dong et al.CVPR 2024 · 15 citations
- IPO: Interpretable Prompt Optimization for Vision-Language ModelsYingjun Du, Wenfang Sun, Cees SnoekNeurIPS 2024 · 15 citations
- ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language ModelsJuntian Zhang, Song Jin, Chuanqi Cheng, Yuhan Liu et al.ICLR 2026 · 7 citations
- A Systematic Survey of Automatic Prompt Optimization TechniquesKiran Ramnath, Kang Zhou, Sheng Guan, Soumya Smruti Mishra et al.EMNLP 2025 · 5 citations
Builds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
Related papers
- Black-Box Tuning for Language-Model-as-a-ServiceTianxiang Sun, Yunfan Shao, Hong Qian, Xuanjing Huang et al.ICML 2022 · 343 citations
- FDPT: Federated Discrete Prompt Tuning for Black-Box Visual-Language ModelsJiaqi Wu, Simin Chen, Jing Tang, Yuzhe Yang et al.ICCV 2025 · 1 citation
- Modality-Agnostic Zeroth-Order LoRA Fine-Tuning for Black-Box Prompt OptimizationXingchen Li, Jia Zhang, Tianxing Man, Wenkang Wang et al.KDD 2026
- InstructZero: Efficient Instruction Optimization for Black-Box Large Language ModelsLichang Chen, Jiuhai Chen, Tom Goldstein, Heng Huang et al.ICML 2024 · 64 citations
- ProAPO: Progressively Automatic Prompt Optimization for Visual ClassificationXiangyan Qu, Gaopeng Gou, Jiamin Zhuang, Jing Yu et al.CVPR 2025
