Bridging Environments and Language with Rendering Functions and Vision-Language Models
Théo Cachet, Christopher R. Dance, Olivier Sigaud
Abstract
Vision-language models (VLMs) have tremendous potential for grounding language, and thus enabling language-conditioned agents (LCAs) to perform diverse tasks specified with text. This has motivated the study of LCAs based on reinforcement learning (RL) with rewards given by rendering images of an environment and evaluating those images with VLMs. If single-task RL is employed, such approaches are limited by the cost and time required to train a policy for each new task. Multi-task RL (MTRL) is a natural alternative, but requires a carefully designed corpus of training tasks and does not always generalize reliably to new tasks. Therefore, this paper introduces a novel decomposition of the problem of building an LCA: first find an environment configuration that has a high VLM score for text describing a task; then use a (pretrained) goalconditioned policy to reach that configuration. We also explore several enhancements to the speed and quality of VLM-based LCAs, notably, the use of distilled models, and the evaluation of configurations from multiple viewpoints to resolve the ambiguities inherent in a single 2D view. We demonstrate our approach on the Humanoid environment, showing that it results in LCAs that outperform MTRL baselines in zero-shot generalization, without requiring any textual task descriptions or other forms of environment-specific annotation during training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aabfa17d-c270-4f65-aab7-0b877972be25Cited by top-tier papers2
- The "Law" of the Unconscious Contrastive Learner: Probabilistic Alignment of Unpaired ModalitiesYongwei Che, Benjamin EysenbachICLR 2025
- Parameter-Masked Decoupled Optimization for Cross-Domain Class-Incremental LearningZiqi Gu, Chunyan Xu, Yangguang Liu, Wenxuan Fang et al.ICML 2026
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez et al.NeurIPS 2023 · 843 citations
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 685 citations
Related papers
- Vision-Language Models are Zero-Shot Reward Models for Reinforcement LearningJuan Rocamonde, Victoriano Montesinos, Elvis Nava, Ethan Perez et al.ICLR 2024 · 154 citations
- Distilling Internet-Scale Vision-Language Models into Embodied AgentsTheodore R. Sumers, Kenneth Marino, Arun Ahuja, Rob Fergus et al.ICML 2023 · 36 citations
- Zero-Shot Reward Specification via Grounded Natural LanguageParsa Mahmoudieh, Deepak Pathak, Trevor DarrellICML 2022 · 69 citations
- GoalLadder: Incremental Goal Discovery with Vision-Language ModelsAlexey Zakharov, Shimon WhitesonNeurIPS 2025 · 4 citations
- MVR: Multi-view Video Reward Shaping for Reinforcement LearningLirui Luo, Guoxi Zhang, Hongming Xu, Yaodong Yang et al.ICLR 2026
