LIV: Language-Image Representations and Rewards for Robotic Control
Yecheng Jason Ma, Vikash Kumar, Amy Zhang, Osbert Bastani, Dinesh Jayaraman
Abstract
We present Language-Image Value learning (LIV), a unified objective for vision-language representation and reward learning from action-free videos with text annotations. Exploiting a novel connection between dual reinforcement learning and mutual information contrastive learning, the LIV objective trains a multi-modal representation that implicitly encodes a universal value function for tasks specified as language or image goals. We use LIV to pre-train the first control-centric vision-language representation from large human video datasets such as EpicKitchen. Given only a language or image goal, the pre-trained LIV model can assign dense rewards to each frame in videos of unseen robots or humans attempting that task in unseen environments. Further, when some target domain-specific data is available, the same objective can be used to fine-tune and improve LIV and even other pre-trained representations for robotic control and reward specification in that domain. In our experiments on several simulated and real-world robot environments, LIV models consistently outperform the best prior input state representations for imitation learning, as well as reward specification methods for policy synthesis. Our results validate the advantages of joint vision-language representation and reward learning within the unified, compact LIV framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 12dfa40c-4aea-411d-9f00-6b173dfcdfceCited by top-tier papers48
- Motif: Intrinsic Motivation from Artificial Intelligence FeedbackMartin Klissarov, Pierluca D'Oro, Shagun Sodhani, Roberta Raileanu et al.ICLR 2024 · 97 citations
- Closed-Loop Visuomotor Control with Generative Expectation for Robotic ManipulationQingwen Bu, Jia Zeng, Li Chen, Yanchao Yang et al.NeurIPS 2024 · 80 citations
- SARM: Stage-Aware Reward Modeling for Long Horizon Robot ManipulationQianzhong Chen, Justin Yu, Mac Schwager, Pieter Abbeel et al.ICLR 2026 · 50 citations
- Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-MakingVivek Myers, Chongyi Zheng, Anca D. Dragan, Sergey Levine et al.ICML 2024 · 38 citations
- Learning an Actionable Discrete Diffusion Policy via Large-Scale Actionless Video Pre-TrainingHaoran He, Chenjia Bai, Ling Pan, Weinan Zhang et al.NeurIPS 2024 · 38 citations
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
Related papers
- VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-TrainingYecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani et al.ICLR 2023 · 35 citations
- EC2: Emergent Communication for Embodied ControlYao Mu, Shunyu Yao, Mingyu Ding, Ping Luo et al.CVPR 2023
- Progressor: A Perceptually Guided Reward Estimator with Self-Supervised Online RefinementTewodros W. Ayalew, Xiao Zhang, Kevin Yuanbo Wu, Tianchong Jiang et al.ICCV 2025 · 13 citations
- MVR: Multi-view Video Reward Shaping for Reinforcement LearningLirui Luo, Guoxi Zhang, Hongming Xu, Yaodong Yang et al.ICLR 2026
- Provable Ordering and Continuity in Vision-Language Pretraining for Generalizable Embodied AgentsZhizhen Zhang, Lei Zhu, Zhen Fang, Zi Huang et al.NeurIPS 2025 · 5 citations
