PEAC: Unsupervised Pre-training for Cross-Embodiment Reinforcement Learning
Chengyang Ying, Zhongkai Hao, Xinning Zhou, Xuezhou Xu, Hang Su, Xingxing Zhang, Jun Zhu
Abstract
Designing generalizable agents capable of adapting to diverse embodiments has achieved significant attention in Reinforcement Learning (RL), which is critical for deploying RL agents in various real-world applications. Previous Cross-Embodiment RL approaches have focused on transferring knowledge across embodiments within specific tasks. These methods often result in knowledge tightly coupled with those tasks and fail to adequately capture the distinct characteristics of different embodiments. To address this limitation, we introduce the notion of Cross-Embodiment Unsupervised RL (CEURL), which leverages unsupervised learning to enable agents to acquire embodiment-aware and task-agnostic knowledge through online interactions within reward-free environments. We formulate CEURL as a novel Controlled Embodiment Markov Decision Process (CE-MDP) and systematically analyze CEURL's pre-training objectives under CE-MDP. Based on these analyses, we develop a novel algorithm Pre-trained Embodiment-Aware Control (PEAC) for handling CEURL, incorporating an intrinsic reward function specifically designed for cross-embodiment pre-training. PEAC not only provides an intuitive optimization strategy for cross-embodiment pre-training but also can integrate flexibly with existing unsupervised RL methods, facilitating cross-embodiment exploration and skill discovery. Extensive experiments in both simulated (e.g., DMC and Robosuite) and real-world environments (e.g., legged locomotion) demonstrate that PEAC significantly improves adaptation performance and cross-embodiment generalization, demonstrating its effectiveness in overcoming the unique challenges of CEURL. The project page and code are in https://yingchengyang.github.io/ceurl.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Exploratory Diffusion Model for Unsupervised Reinforcement LearningChengyang Ying, Huayu Chen, Xinning Zhou, Zhongkai Hao et al.ICLR 2026 · 4 citations
- GraphMimic: Graph-to-Graphs Generative Modeling from Videos for Policy LearningGuangyan Chen, Te Cui, Meiling Wang, Chengcai Yang et al.CVPR 2025
- Vintix: Action Model via In-Context Reinforcement LearningAndrei Polubarov, Nikita Lyubaykin, Alexander Derevyagin, Ilya Zisman et al.ICML 2025
Builds on40
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 911 citations
Related papers
- Cross-Embodiment Offline Reinforcement Learning for Heterogeneous Robot DatasetsHaruki Abe, Takayuki Osa, Yusuke Mukuta, Tatsuya HaradaICLR 2026 · 1 citation
- Behavior From the Void: Unsupervised Active Pre-TrainingHao Liu, Pieter AbbeelNeurIPS 2021 · 258 citations
- Improving Compositional Generalization in Cross-Embodiment Learning via Mixture of Disentangled PrototypesRen Wang, Xin Wang, Tongtong Feng, Xinyue Gong et al.ACM MM 2025 · 2 citations
- Mastering the Unsupervised Reinforcement Learning Benchmark from PixelsSai Rajeswar, Pietro Mazzaglia, Tim Verbelen, Alexandre Piché et al.ICML 2023 · 30 citations
- Self-Supervised Reinforcement Learning that Transfers using Random FeaturesBoyuan Chen, Chuning Zhu, Pulkit Agrawal, Kaiqing Zhang et al.NeurIPS 2023 · 16 citations
