The Unsurprising Effectiveness of Pre-Trained Vision Models for Control
Simone Parisi, Aravind Rajeswaran, Senthil Purushwalkam, Abhinav Gupta
摘要
Recent years have seen the emergence of pretrained representations as a powerful abstraction for AI applications in computer vision, natural language, and speech. However, policy learning for control is still dominated by a tabularasa learning paradigm, with visuo-motor policies often trained from scratch using data from deployment environments. In this context, we revisit and study the role of pre-trained visual representations for control, and in particular representations trained on large-scale computer vision datasets. Through extensive empirical evaluation in diverse control domains (Habitat, Deep-Mind Control, Adroit, Franka Kitchen), we isolate and study the importance of different representation training methods, data augmentations, and feature hierarchies. Overall, we find that pre-trained visual representations can be competitive or even better than ground-truth state representations to train control policies. This is in spite of using only out-of-domain data from standard vision datasets, without any in-domain data from the deployment environments. Source code and more at https://sites.google. com/view/pvr-control .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper70
- Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Yecheng Jason Ma 等NeurIPS 2023 · 被引用 336 次
- Unleashing Large-Scale Video Generative Pre-training for Visual Robot ManipulationHongtao Wu, Ya Jing, Chilam Cheang, Guangzeng Chen 等ICLR 2024 · 被引用 309 次
- LIV: Language-Image Representations and Rewards for Robotic ControlYecheng Jason Ma, Vikash Kumar, Amy Zhang, Osbert Bastani 等ICML 2023 · 被引用 212 次
- Pre-Trained Image Encoder for Generalizable Visual Reinforcement LearningZhecheng Yuan, Zhengrong Xue, Bo Yuan, Xueqian Wang 等NeurIPS 2022 · 被引用 112 次
- Multi-View Masked World Models for Visual Robotic ManipulationYounggyo Seo, Junsu Kim, Stephen James, Kimin Lee 等ICML 2023 · 被引用 99 次
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
相关 Paper
- On Pre-Training for Visuo-Motor Control: Revisiting a Learning-from-Scratch BaselineNicklas Hansen, Zhecheng Yuan, Yanjie Ze, Tongzhou Mu 等ICML 2023 · 被引用 78 次
- For Pre-Trained Vision Models in Motor Control, Not All Policy Learning Methods are Created EqualYingdong Hu, Renhao Wang, Li Erran Li, Yang GaoICML 2023 · 被引用 28 次
- Pre-trained Text-to-Image Diffusion Models Are Versatile Representation Learners for ControlGunshi Gupta, Karmesh Yadav, Yarin Gal, Dhruv Batra 等NeurIPS 2024 · 被引用 17 次
- Exploring Conditions for Diffusion Models in Robotic ControlHeeseong Shin, Byeongho Heo, Dongyoon Han, Seungryong Kim 等CVPR 2026
- Behavior From the Void: Unsupervised Active Pre-TrainingHao Liu, Pieter AbbeelNeurIPS 2021 · 被引用 258 次
