Improving Compositional Generalization in Cross-Embodiment Learning via Mixture of Disentangled Prototypes
Ren Wang, Xin Wang, Tongtong Feng, Xinyue Gong, Guangyao Li, Yu-Wei Zhan, Qing Li, Wenwu Zhu
Abstract
Cross-Embodiment Learning (CEL) aims to train a generalist policy model by integrating large-scale compositional interactions of heterogeneous agents and environments. However, the inherent conflict between the unbounded space of agent-environment combinations and a single unified policy model hinders generalization to unseen combinations. To address this challenge, we propose a novel Mixture of Disentangled Prototypes (MoDP) method to improve the compositional generalization in CEL. The key idea is to introduce a finite prototype space that bridges the gap between unbounded agent-environment combinations and a single policy model. Specifically, we design a dual-headed autoencoder and a compositional reconstruction loss to disentangle agent and environment features from interaction data, and map them into respective prototype spaces. We then introduce a connection-sensitivity-based pruning method to extract sub-networks from the pre-trained policy model, forming policy prototypes associated with specific agent-environment prototype pairs. Finally, a parameter-free routing mechanism adaptively integrates relevant policy prototypes for each input composition. Experiments in both standard and compositional settings demonstrate the effectiveness of our MoDP in enhancing the generalization capability of pre-trained policies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- U2UData+: A Scalable Swarm UAVs Autonomous Flight Dataset for Embodied Long-horizon TasksTongtong Feng, Xin Wang, Feilin Han, Leping Zhang et al.AAAI 2026 · 4 citations
- ModularAgent: A Task-Aware Modular Framework for Joint Optimization of Multimodal Large Language Models and World ModelsYu-Wei Zhan, Xin Wang, Pengzhe Mao, Tongtong Feng et al.CVPR 2026
Builds on21
- Object-Centric Learning with Slot AttentionFrancesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran et al.NeurIPS 2020 · 1,275 citations
- Contrastive Learning of Structured World ModelsThomas N. Kipf, Elise van der Pol, Max WellingICLR 2020 · 322 citations
- Disentangled Self-Supervision in Sequential RecommendersJianxin Ma, Chang Zhou, Hongxia Yang, Peng Cui et al.KDD 2020 · 223 citations
- One Policy to Control Them All: Shared Modular Policies for Agent-Agnostic ControlWenlong Huang, Igor Mordatch, Deepak PathakICML 2020 · 214 citations
- Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained TransformersLirui Wang, Xinlei Chen, Jialiang Zhao, Kaiming HeNeurIPS 2024 · 208 citations
Related papers
- Mixture of Meta-Policies for Cross-Environment Meta-Reinforcement LearningXinyu Liu, Qingyu Zeng, Chenwei Tang, Jiancheng LvKDD 2026
- Cross-Embodiment Offline Reinforcement Learning for Heterogeneous Robot DatasetsHaruki Abe, Takayuki Osa, Yusuke Mukuta, Tatsuya HaradaICLR 2026 · 1 citation
- PEAC: Unsupervised Pre-training for Cross-Embodiment Reinforcement LearningChengyang Ying, Zhongkai Hao, Xinning Zhou, Xuezhou Xu et al.NeurIPS 2024 · 14 citations
- Rethinking Missing Modality Learning from a Decoding PerspectiveTao Jin, Xize Cheng, Linjun Li, Wang Lin et al.ACM MM 2023 · 10 citations
- Improving Generalization in Reinforcement Learning with Mixture RegularizationKaixin Wang, Bingyi Kang, Jie Shao, Jiashi FengNeurIPS 2020 · 143 citations
