Learning Vision-Guided Quadrupedal Locomotion End-to-End with Cross-Modal Transformers
Ruihan Yang, Minghao Zhang, Nicklas Hansen, Huazhe Xu, Xiaolong Wang
Abstract
We propose to address quadrupedal locomotion tasks using Reinforcement Learning (RL) with a Transformer-based model that learns to combine proprioceptive information and high-dimensional depth sensor inputs. While learning-based locomotion has made great advances using RL, most methods still rely on domain randomization for training blind agents that generalize to challenging terrains. Our key insight is that proprioceptive states only offer contact measurements for immediate reaction, whereas an agent equipped with visual sensory observations can learn to proactively maneuver environments with obstacles and uneven terrain by anticipating changes in the environment many steps ahead. In this paper, we introduce LocoTransformer, an end-to-end RL method that leverages both proprioceptive states and visual observations for locomotion control. We evaluate our method in challenging simulated environments with different obstacles and uneven terrain. We transfer our learned policy from simulation to a real robot by running it indoors and in the wild with unseen obstacles and terrain. Our method not only significantly improves over baselines, but also achieves far better generalization performance, especially when transferred to the real robot. Our project page with videos is at https://rchalyang.github.io/LocoTransformer/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a0423ce-3d17-4d07-a435-d5ac72e1972cCited by top-tier papers13
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang et al.NeurIPS 2022 · 1,291 citations
- Hybrid Internal Model: Learning Agile Legged Locomotion with Simulated Robot ResponseJunfeng Long, Zirui Wang, Quanyi Li, Liu Cao et al.ICLR 2024 · 66 citations
- Coupling Vision and Proprioception for Navigation of Legged RobotsZipeng Fu, Ashish Kumar, Ananye Agarwal, Haozhi Qi et al.CVPR 2022 · 36 citations
- Uni-O4: Unifying Online and Offline Deep Reinforcement Learning with Multi-Step On-Policy OptimizationKun Lei, Zhengmao He, Chenhao Lu, Kaizhe Hu et al.ICLR 2024 · 31 citations
- MoVie: Visual Model-Based Policy Adaptation for View GeneralizationSizhe Yang, Yanjie Ze, Huazhe XuNeurIPS 2023 · 29 citations
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
Related papers
- Continuous Spatiotemporal TransformerAntonio Henrique de Oliveira Fonseca, Emanuele Zappala, Josue Ortega Caro, David van DijkICML 2023 · 2 citations
- CORN: Contact-based Object Representation for Nonprehensile Manipulation of General Unseen ObjectsYoonyoung Cho, Junhyek Han, Yoontae Cho, Beomjoon KimICLR 2024 · 20 citations
- Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained TransformersLirui Wang, Xinlei Chen, Jialiang Zhao, Kaiming HeNeurIPS 2024 · 208 citations
- RRL: Resnet as representation for Reinforcement LearningRutav M. Shah, Vikash KumarICML 2021 · 129 citations
- Skill Transformer: A Monolithic Policy for Mobile ManipulationXiaoyu Huang, Dhruv Batra, Akshara Rai, Andrew SzotICCV 2023 · 34 citations
