Multi-Environment Pretraining Enables Transfer to Action Limited Datasets
David Venuto, Sherry Yang, Pieter Abbeel, Doina Precup, Igor Mordatch, Ofir Nachum
Abstract
Using massive datasets to train large-scale models has emerged as a dominant approach for broad generalization in natural language and vision applications. In reinforcement learning, however, a key challenge is that available data of sequential decision making is often not annotated with actions - for example, videos of game-play are much more available than sequences of frames paired with their logged game controls. We propose to circumvent this challenge by combining large but sparsely-annotated datasets from a target environment of interest with fully-annotated datasets from various other source environments. Our method, Action Limited PreTraining (ALPT), leverages the generalization capabilities of inverse dynamics modelling (IDM) to label missing action data in the target environment. We show that utilizing even one additional environment dataset of labelled data during IDM pretraining gives rise to substantial improvements in generating action labels for unannotated sequences. We evaluate our method on benchmark game-playing environments and show that we can significantly improve game performance and generalization capability compared to other approaches, using annotated datasets equivalent to only minutes of gameplay. Highlighting the power of IDM, we show that these benefits remain even when target and source environments share no common actions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Inverse Dynamics Pretraining Learns Good Representations for Multitask ImitationDavid Brandfonbrener, Ofir Nachum, Joan BrunaNeurIPS 2023 · 38 citations
- Mastering Robot Manipulation with Multimodal Prompts through Pretraining and Multi-task Fine-tuningJiachen Li, Qiaozi Gao, Michael Johnston, Xiaofeng Gao et al.ICML 2024 · 18 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
Related papers
- Learning to Act without ActionsDominik Schmidt, Minqi JiangICLR 2024 · 98 citations
- On the Sample Efficiency of Inverse Dynamics Models for Semi-Supervised Imitation LearningSacha Morin, Moonsub Byeon, Alexia Jolicoeur-Martineau, Sebastien LachapelleICML 2026 · 1 citation
- Learning Video-Conditioned Policy on Unlabelled Data with Joint Embedding Predictive TransformerHao Luo, Zongqing LuICLR 2025
- Semi-Supervised Offline Reinforcement Learning with Action-Free TrajectoriesQinqing Zheng, Mikael Henaff, Brandon Amos, Aditya GroverICML 2023 · 29 citations
- Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the WildHao Luo, Ye Wang, Wanpeng Zhang, Haoqi Yuan et al.CVPR 2026 · 15 citations
