MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-Scenarios
Xuantang Xiong, Ni Mu, Runpeng Xie, Senhao Yang, Yaqing Wang, Lexiang Wang, Yao Luan, Siyuan Li, Shuang Xu, Yiqin Yang, Bo Xu
Abstract
Model-based reinforcement learning (MBRL) is a crucial approach to enhance the generalization capabilities and improve the sample efficiency of RL algorithms. However, current MBRL methods focus primarily on building world models for single tasks and rarely address generalization across different scenarios. Building on the insight that dynamics within the same simulation engine share inherent properties, we attempt to construct a unified world model capable of generalizing across different scenarios, named Meta-Regularized Contextual World-Model (MrCoM). This method first decomposes the latent state space into various components based on the dynamic characteristics, thereby enhancing the accuracy of world-model prediction. Further, MrCoM adopts meta-state regularization to extract unified representation of scenario-relevant information, and meta-value regularization to align world-model optimization with policy learning across diverse scenario objectives. We theoretically analyze the generalization error upper bound of MrCoM in multi-scenario settings. We systematically evaluate our algorithm's generalization ability across diverse scenarios, demonstrating significantly better performance than previous state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d8be3a3-691e-4170-a157-cd3b48796bc7Builds on5
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- Context-aware Dynamics Model for Generalization in Model-Based Reinforcement LearningKimin Lee, Younggyo Seo, Seunghyun Lee, Honglak Lee et al.ICML 2020 · 158 citations
- Proper Value EquivalenceChristopher Grimm, André Barreto, Gregory Farquhar, David Silver et al.NeurIPS 2021 · 49 citations
- Value Gradient weighted Model-Based Reinforcement LearningClaas Voelcker, Victor Liao, Animesh Garg, Amir-massoud FarahmandICLR 2022 · 37 citations
- Live in the Moment: Learning Dynamics Model Adapted to Evolving PolicyXiyao Wang, Wichayaporn Wongkamjan, Ruonan Jia, Furong HuangICML 2023 · 20 citations
Related papers
- Modeling Unseen Environments with Language-guided Composable Causal Components in Reinforcement LearningXinyue Wang, Biwei HuangICLR 2025
- TaskLoom: Weaving Knowledge Across Tasks in World ModelsQingzhang Zeng, Peixi Peng, hang li, Luntong Li et al.ICML 2026
- Pre-training Contextualized World Models with In-the-wild Videos for Reinforcement LearningJialong Wu, Haoyu Ma, Chaoyi Deng, Mingsheng LongNeurIPS 2023 · 55 citations
- Trajectory-wise Multiple Choice Learning for Dynamics Generalization in Reinforcement LearningYounggyo Seo, Kimin Lee, Ignasi Clavera Gilaberte, Thanard Kurutach et al.NeurIPS 2020 · 51 citations
- CATAL: Causally Disentangled Task Representation Learning for Offline Meta-Reinforcement LearningShan Cong, Chao Yu, Xiangyuan LanAAAI 2026
