Become a Proficient Player with Limited Data through Watching Pure Videos
Weirui Ye, Yunsheng Zhang, Pieter Abbeel, Yang Gao
Abstract
Recently, RL has shown its strong ability for visually complex tasks. However, it suffers from the low sample efficiency and poor generalization ability, which prevent RL from being useful in real-world scenarios. Inspired by the huge success of unsupervised pre-training methods on language and vision domains, we propose to improve the sample efficiency via a novel pre-training method for model-based RL. Instead of using pre-recorded agent trajectories that come with their own actions, we consider the setting where the pre-training data are action-free videos, which are more common and available in the real world. We introduce a two-phase training pipeline as follows: for the pre-training phase, we implicitly extract the hidden action embedding from videos and pre-train the visual representation and the environment dynamics network through a novel forward-inverse cycle consistency (FICC) objective based on vector quantization; for down-stream tasks, we finetune with small amount of task data based on the learned models. Our framework can significantly improve the sample efficiency on Atari Games with data of only one hour of game playing. We achieve 118.4% mean human performance and 36.0% median performance with only 50k environment steps, which is 85.6% and 65.1% better than the scratch EfficientZero model. We believe such pre-training approach can provide an option for solving real-world RL problems. The code is available at https://github.com/YeWR/FICC.git . Latent action Cycle consistency Inverse Forward
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 346f3a08-c3eb-4202-98f1-4a02ee69b426Cited by top-tier papers29
- Learning to Act without ActionsDominik Schmidt, Minqi JiangICLR 2024 · 98 citations
- DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor ControlZichen Jeff Cui, Hengkai Pan, Aadhithya Iyer, Siddhant Haldar et al.NeurIPS 2024 · 61 citations
- Pre-training Contextualized World Models with In-the-wild Videos for Reinforcement LearningJialong Wu, Haoyu Ma, Chaoyi Deng, Mingsheng LongNeurIPS 2023 · 55 citations
- What Do Latent Action Models Actually Learn?Chuheng Zhang, Tim Pearce, Pushi Zhang, Kaixin Wang et al.NeurIPS 2025 · 35 citations
- Unsupervised Behavior Extraction via Random Intent PriorsHao Hu, Yiqin Yang, Jianing Ye, Ziqing Mai et al.NeurIPS 2023 · 15 citations
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
Related papers
- Reinforcement Learning with Action-Free Pre-Training from VideosYounggyo Seo, Kimin Lee, Stephen James, Pieter AbbeelICML 2022 · 150 citations
- PlayVirtual: Augmenting Cycle-Consistent Virtual Trajectories for Reinforcement LearningTao Yu, Cuiling Lan, Wenjun Zeng, Mingxiao Feng et al.NeurIPS 2021 · 64 citations
- Pretraining Representations for Data-Efficient Reinforcement LearningMax Schwarzer, Nitarshan Rajkumar, Michael Noukhovitch, Ankesh Anand et al.NeurIPS 2021 · 151 citations
- The Surprising Ineffectiveness of Pre-Trained Visual Representations for Model-Based Reinforcement LearningMoritz Schneider, Robert Krug, Narunas Vaskevicius, Luigi Palmieri et al.NeurIPS 2024 · 10 citations
- Value-Consistent Representation Learning for Data-Efficient Reinforcement LearningYang Yue, Bingyi Kang, Zhongwen Xu, Gao Huang et al.AAAI 2023 · 19 citations
