Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement
Zaid Khan, Vijay Kumar B. G, Samuel Schulter, Yun Fu, Manmohan Chandraker
摘要
Visual program synthesis is a promising approach to exploit the reasoning abilities of large language models for compositional computer vision tasks. Previous work has used few-shot prompting with frozen LLMs to synthesize visual programs. Training an LLM to write better visual programs is an attractive prospect, but it is unclear how to accomplish this. No dataset of visual programs for training exists, and acquisition of a visual program dataset cannot be easily crowdsourced due to the need for expert annotators. To get around the lack of direct supervision, we explore improving the program synthesis abilities of an LLM using feedback from interactive experience. We propose a method where we exploit existing annotations for a vision-language task to improvise a coarse reward signal for that task, treat the LLM as a policy, and apply reinforced self-training to improve the visual program synthesis ability of the LLM for that task. We describe a series of experiments on object detection, compositional visual question answering, and image-text retrieval, and show that in each case, the selftrained LLM outperforms or performs on par with few-shot frozen LLMs that are an order of magnitude larger. Website: https://zaidkhan.me/ViReP
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- MATA: A Trainable Hierarchical Automaton System for Multi-Agent Visual ReasoningZhixi Cai, Fucai Ke, Kevin Leo, Sukai Huang 等ICLR 2026 · 被引用 3 次
- DWIM: Towards Tool-Aware Visual Reasoning via Discrepancy-Aware Workflow Generation & Instruct-Masking TuningFucai Ke, Vijay Kumar B. G, Xingjian Leng, Zhixi Cai 等ICCV 2025 · 被引用 1 次
- ViUniT: Visual Unit Tests for More Robust Visual ProgrammingArtemis Panagopoulou, Honglu Zhou, Silvio Savarese, Caiming Xiong 等CVPR 2025
- S³-MSD: Large Vision-Language Model for Explainable and Generalizable Multi-modal Sarcasm DetectionZhihong Zhu, Fan Zhang, Yunyan Zhang, Jinghan Sun 等AAAI 2026
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
相关 Paper
- Few-Shot Composition Learning for Image Retrieval with Prompt TuningJunda Wu, Rui Wang, Handong Zhao, Ruiyi Zhang 等AAAI 2023 · 被引用 16 次
- From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data SynthesisChuanqi Cheng, Jian Guan, Wei Wu, Rui YanEMNLP 2024 · 被引用 1 次
- Visual Programming: Compositional visual reasoning without trainingTanmay Gupta, Aniruddha KembhaviCVPR 2023
- Unveiling the Compositional Ability Gap in Vision-Language Reasoning ModelTianle Li, Jihai Zhang, Yongming Rao, Yu ChengNeurIPS 2025 · 被引用 17 次
- STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-TrainingHaiyi Qiu, Minghe Gao, Long Qian, Kaihang Pan 等CVPR 2025
