Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement
Zaid Khan, Vijay Kumar B. G, Samuel Schulter, Yun Fu, Manmohan Chandraker
Abstract
Visual program synthesis is a promising approach to exploit the reasoning abilities of large language models for compositional computer vision tasks. Previous work has used few-shot prompting with frozen LLMs to synthesize visual programs. Training an LLM to write better visual programs is an attractive prospect, but it is unclear how to accomplish this. No dataset of visual programs for training exists, and acquisition of a visual program dataset cannot be easily crowdsourced due to the need for expert annotators. To get around the lack of direct supervision, we explore improving the program synthesis abilities of an LLM using feedback from interactive experience. We propose a method where we exploit existing annotations for a vision-language task to improvise a coarse reward signal for that task, treat the LLM as a policy, and apply reinforced self-training to improve the visual program synthesis ability of the LLM for that task. We describe a series of experiments on object detection, compositional visual question answering, and image-text retrieval, and show that in each case, the selftrained LLM outperforms or performs on par with few-shot frozen LLMs that are an order of magnitude larger. Website: https://zaidkhan.me/ViReP
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9516cdc0-323e-4637-a10a-aaf4018afb3cCited by top-tier papers4
- MATA: A Trainable Hierarchical Automaton System for Multi-Agent Visual ReasoningZhixi Cai, Fucai Ke, Kevin Leo, Sukai Huang et al.ICLR 2026 · 3 citations
- DWIM: Towards Tool-Aware Visual Reasoning via Discrepancy-Aware Workflow Generation & Instruct-Masking TuningFucai Ke, Vijay Kumar B. G, Xingjian Leng, Zhixi Cai et al.ICCV 2025 · 1 citation
- ViUniT: Visual Unit Tests for More Robust Visual ProgrammingArtemis Panagopoulou, Honglu Zhou, Silvio Savarese, Caiming Xiong et al.CVPR 2025
- S³-MSD: Large Vision-Language Model for Explainable and Generalizable Multi-modal Sarcasm DetectionZhihong Zhu, Fan Zhang, Yunyan Zhang, Jinghan Sun et al.AAAI 2026
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
Related papers
- Few-Shot Composition Learning for Image Retrieval with Prompt TuningJunda Wu, Rui Wang, Handong Zhao, Ruiyi Zhang et al.AAAI 2023 · 16 citations
- From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data SynthesisChuanqi Cheng, Jian Guan, Wei Wu, Rui YanEMNLP 2024 · 1 citation
- Visual Programming: Compositional visual reasoning without trainingTanmay Gupta, Aniruddha KembhaviCVPR 2023
- Unveiling the Compositional Ability Gap in Vision-Language Reasoning ModelTianle Li, Jihai Zhang, Yongming Rao, Yu ChengNeurIPS 2025 · 17 citations
- STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-TrainingHaiyi Qiu, Minghe Gao, Long Qian, Kaihang Pan et al.CVPR 2025
