Masked Autoencoding for Scalable and Generalizable Decision Making
Fangchen Liu, Hao Liu, Aditya Grover, Pieter Abbeel
摘要
We are interested in learning scalable agents for reinforcement learning that can learn from large-scale, diverse sequential data similar to current large vision and language models. To this end, this paper presents masked decision prediction (MaskDP), a simple and scalable self-supervised pretraining method for reinforcement learning (RL) and behavioral cloning (BC). In our MaskDP approach, we employ a masked autoencoder (MAE) to state-action trajectories, wherein we randomly mask state and action tokens and reconstruct the missing data. By doing so, the model is required to infer masked-out states and actions and extract information about dynamics. We find that masking different proportions of the input sequence significantly helps with learning a better model that generalizes well to multiple downstream tasks. In our empirical study, we find that a MaskDP model gains the capability of zero-shot transfer to new BC tasks, such as single and multiple goal reaching, and it can zero-shot infer skills from a few example transitions. In addition, MaskDP transfers well to offline RL and shows promising scaling behavior w.r.t. to model size. It is amenable to data-efficient finetuning, achieving competitive results with prior methods based on autoregressive pretraining 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Unleashing Large-Scale Video Generative Pre-training for Visual Robot ManipulationHongtao Wu, Ya Jing, Chilam Cheang, Guangzeng Chen 等ICLR 2024 · 被引用 309 次
- Diffusion Models for Black-Box OptimizationSiddarth Krishnamoorthy, Satvik Mehul Mashkaria, Aditya GroverICML 2023 · 被引用 94 次
- VIMA: Robot Manipulation with Multimodal PromptsYunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang 等ICML 2023 · 被引用 80 次
- Foundation Policies with Hilbert RepresentationsSeohong Park, Tobias Kreiman, Sergey LevineICML 2024 · 被引用 72 次
- Masked Trajectory Models for Prediction, Representation, and ControlPhilipp Wu, Arjun Majumdar, Kevin Stone, Yixin Lin 等ICML 2023 · 被引用 57 次
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
相关 Paper
- Self-Supervised Reinforcement Learning that Transfers using Random FeaturesBoyuan Chen, Chuning Zhu, Pulkit Agrawal, Kaiqing Zhang 等NeurIPS 2023 · 被引用 16 次
- RePreM: Representation Pre-training with Masked Model for Reinforcement LearningYuanying Cai, Chuheng Zhang, Wei Shen, Xuyun Zhang 等AAAI 2023 · 被引用 7 次
- Uni[MASK]: Unified Inference in Sequential Decision ProblemsMicah Carroll, Orr Paradise, Jessy Lin, Raluca Georgescu 等NeurIPS 2022 · 被引用 29 次
- Learning Versatile Skills with Curriculum MaskingYao Tang, Zhihui Xie, Zichuan Lin, Deheng Ye 等NeurIPS 2024 · 被引用 6 次
- SMART: Self-supervised Multi-task pretrAining with contRol TransformersYanchao Sun, Shuang Ma, Ratnesh Madaan, Rogerio Bonatti 等ICLR 2023 · 被引用 4 次
