Learning Video-Conditioned Policy on Unlabelled Data with Joint Embedding Predictive Transformer
Hao Luo, Zongqing Lu
摘要
The video-conditioned policy takes prompt videos of the desired tasks as a condition and is regarded for its prospective generalizability. Despite its promise, training a video-conditioned policy is non-trivial due to the need for abundant demonstrations. In some tasks, the expert rollouts are merely available as videos, and costly and time-consuming efforts are required to annotate action labels. To address this, we explore training video-conditioned policy on a mixture of demonstrations and unlabeled expert videos to reduce reliance on extensive manual annotation. We introduce the Joint Embedding Predictive Transformer (JEPT) to learn a videoconditioned policy through sequence modeling. JEPT is designed to jointly learn visual transition prediction and inverse dynamics. The visual transition is captured from both demonstrations and expert videos, on the basis of which the inverse dynamics learned from demonstrations is generalizable to the tasks without action labels. Experiments on a series of simulated visual control tasks evaluate that JEPT can effectively leverage the mixture dataset to learn a generalizable policy. JEPT outperforms baselines in the tasks without action-labeled data and unseen tasks. We also experimentally reveal the potential of JEPT as a simple visual priors injection approach to enhance the video-conditioned policy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Disentangled Robot Learning via Separate Forward and Inverse Dynamics PretrainingWenyao Zhang, Bozhou Zhang, Zekun Qi, Wenjun Zeng 等ICLR 2026 · 被引用 18 次
- Human2Robot: Learning Robot Actions from Paired Human-Robot VideosSicheng Xie, Haidong Cao, Zejia Weng, Zhen Xing 等AAAI 2026 · 被引用 15 次
- Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the WildHao Luo, Ye Wang, Wanpeng Zhang, Haoqi Yuan 等CVPR 2026 · 被引用 15 次
- Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action GenerationXiangkai Ma, Lekai Xing, Han Zhang, Wenzhong Li 等CVPR 2026 · 被引用 11 次
- RLZero: Direct Policy Inference from Language Without In-Domain SupervisionHarshit Sikchi, Siddhant Agarwal, Pranaya Jajoo, Samyak Parajuli 等NeurIPS 2025 · 被引用 8 次
它引用的顶会 Paper23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Perceiver IO: A General Architecture for Structured Inputs & OutputsAndrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch 等ICLR 2022 · 被引用 797 次
- Learning Universal Policies via Text-Guided Video GenerationYilun Du, Sherry Yang, Bo Dai, Hanjun Dai 等NeurIPS 2023 · 被引用 742 次
相关 Paper
- Video Prediction Models as Rewards for Reinforcement LearningAlejandro Escontrela, Ademi Adeniji, Wilson Yan, Ajay Jain 等NeurIPS 2023 · 被引用 117 次
- Solving New Tasks by Adapting Internet Video KnowledgeCalvin Luo, Zilai Zeng, Yilun Du, Chen SunICLR 2025
- Motus: A Unified Latent Action World ModelHongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang 等CVPR 2026 · 被引用 271 次
- From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot ManipulationYajie Li, Bozhou Zhang, Chun Gu, Zipei Ma 等ICML 2026 · 被引用 2 次
- From Play to Policy: Conditional Behavior Generation from Uncurated Robot DataZichen Jeff Cui, Yibin Wang, Nur Muhammad (Mahi) Shafiullah, Lerrel PintoICLR 2023 · 被引用 6 次
