Reward-Free Curricula for Training Robust World Models
Marc Rigter, Minqi Jiang, Ingmar Posner
摘要
There has been a recent surge of interest in developing generally-capable agents that can adapt to new tasks without additional training in the environment. Learning world models from reward-free exploration is a promising approach, and enables policies to be trained using imagined experience for new tasks. However, achieving a general agent requires robustness across different environments. In this work, we address the novel problem of generating curricula in the reward-free setting to train robust world models. We consider robustness in terms of minimax regret over all environment instantiations and show that the minimax regret can be connected to minimising the maximum error in the world model across environment instances. This result informs our algorithm, WAKER: Weighted Acquisition of Knowledge across Environments for Robustness. WAKER selects environments for data collection based on the estimated error of the world model for each environment. Our experiments demonstrate that WAKER outperforms several baselines, resulting in improved robustness, efficiency, and generalisation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- No Regrets: Investigating and Improving Regret Approximations for Curriculum DiscoveryAlexander Rutherford, Michael Beukman, Timon Willi, Bruno Lacerda 等NeurIPS 2024 · 被引用 38 次
- Leveraging Separated World Model for Exploration in Visually Distracted EnvironmentsKaichen Huang, Shenghua Wan, Minghao Shao, Hai-Hang Sun 等NeurIPS 2024 · 被引用 5 次
- Improving Regret Approximation for Unsupervised Dynamic Environment GenerationHarry Mead, Bruno Lacerda, Jakob N. Foerster, Nick HawesNeurIPS 2025 · 被引用 1 次
- TRACED: Transition-aware Regret Approximation with Co-learnability for Environment DesignGeonwoo Cho, Jaegyun Im, Jihwan Lee, Hojun Yi 等ICLR 2026 · 被引用 1 次
- Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured ModelingFan Feng, Yujia Zheng, Minghao Fu, Yongqiang Chen 等ICML 2026
它引用的顶会 Paper21
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 被引用 870 次
相关 Paper
- Imagined AutocurriculaAhmet Hamdi Güzel, Matthew Thomas Jackson, Jarek Liesen, Tim Rocktäschel 等NeurIPS 2025 · 被引用 2 次
- Adversarial Environment Design via Regret-Guided Diffusion ModelsHojun Chung, Junseo Lee, Minsoo Kim, Dohyeong Kim 等NeurIPS 2024 · 被引用 11 次
- Emergent Complexity and Zero-shot Transfer via Unsupervised Environment DesignMichael Dennis, Natasha Jaques, Eugene Vinitsky, Alexandre M. Bayen 等NeurIPS 2020 · 被引用 362 次
- General agents need world modelsJonathan Richens, Tom Everitt, David AbelICML 2025
- Learning Temporally AbstractWorld Models without Online ExperimentationBenjamin Freed, Siddarth Venkatraman, Guillaume Adrien Sartoretti, Jeff Schneider 等ICML 2023 · 被引用 7 次
