For Pre-Trained Vision Models in Motor Control, Not All Policy Learning Methods are Created Equal
Yingdong Hu, Renhao Wang, Li Erran Li, Yang Gao
摘要
In recent years, increasing attention has been directed to leveraging pre-trained vision models for motor control. While existing works mainly emphasize the importance of this pre-training phase, the arguably equally important role played by downstream policy learning during controlspecific fine-tuning is often neglected. It thus remains unclear if pre-trained vision models are consistent in their effectiveness under different control policies. To bridge this gap in understanding, we conduct a comprehensive study on 14 pre-trained vision models using 3 distinct classes of policy learning methods, including reinforcement learning (RL), imitation learning through behavior cloning (BC), and imitation learning with a visual reward function (VRF). Our study yields a series of intriguing results, including the discovery that the effectiveness of pre-training is highly dependent on the choice of the downstream policy learning algorithm. We show that conventionally accepted evaluation based on RL methods is highly variable and therefore unreliable, and further advocate for using more robust methods like VRF and BC. To facilitate more universal evaluations of pre-trained models and their policy learning methods in the future, we also release a benchmark of 21 tasks across 3 different environments alongside our work. Source code and more details can be found at https://yingdong-hu.github. io/PVM-control/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Revealing Vision-Language Integration in the Brain with Multimodal NetworksVighnesh Subramaniam, Colin Conwell, Christopher Wang, Gabriel Kreiman 等ICML 2024 · 被引用 19 次
- Learning Interactive World Model for Object-Centric Reinforcement LearningFan Feng, Phillip Lippe, Sara MagliacaneNeurIPS 2025 · 被引用 13 次
- The Surprising Ineffectiveness of Pre-Trained Visual Representations for Model-Based Reinforcement LearningMoritz Schneider, Robert Krug, Narunas Vaskevicius, Luigi Palmieri 等NeurIPS 2024 · 被引用 10 次
- Efficient Reinforcement Learning Through Adaptively Pretrained Visual EncoderYuhan Zhang, Guoqing Ma, Guangfu Hao, Liangxuan Guo 等AAAI 2025 · 被引用 3 次
- Capturing Visual Environment Structure Correlates with Control PerformanceJiahua Dong, Yunze Man, Pavel Tokmakov, Yu-Xiong WangICLR 2026 · 被引用 2 次
它引用的顶会 Paper38
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
相关 Paper
- The Unsurprising Effectiveness of Pre-Trained Vision Models for ControlSimone Parisi, Aravind Rajeswaran, Senthil Purushwalkam, Abhinav GuptaICML 2022 · 被引用 233 次
- On Pre-Training for Visuo-Motor Control: Revisiting a Learning-from-Scratch BaselineNicklas Hansen, Zhecheng Yuan, Yanjie Ze, Tongzhou Mu 等ICML 2023 · 被引用 78 次
- Vision-Language Foundation Models as Effective Robot ImitatorsXinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu 等ICLR 2024 · 被引用 375 次
- Posterior Behavioral Cloning: Pretraining BC Policies for Efficient RL FinetuningAndrew Wagenmaker, Perry Dong, Raymond Tsao, Chelsea Finn 等ICML 2026 · 被引用 10 次
- Where are we in the search for an Artificial Visual Cortex for Embodied Intelligence?Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Yecheng Jason Ma 等NeurIPS 2023 · 被引用 336 次
