Learning World Models for Unconstrained Goal Navigation
Yuanlin Duan, Wensen Mao, He Zhu
摘要
Learning world models offers a promising avenue for goal-conditioned reinforcement learning with sparse rewards. By allowing agents to plan actions or exploratory goals without direct interaction with the environment, world models enhance exploration efficiency. The quality of a world model hinges on the richness of data stored in the agent's replay buffer, with expectations of reasonable generalization across the state space surrounding recorded trajectories. However, challenges arise in generalizing learned world models to state transitions backward along recorded trajectories or between states across different trajectories, hindering their ability to accurately model real-world dynamics. To address these challenges, we introduce a novel goal-directed exploration algorithm, MUN (short for"World Models for Unconstrained Goal Navigation"). This algorithm is capable of modeling state transitions between arbitrary subgoal states in the replay buffer, thereby facilitating the learning of policies to navigate between any"key"states. Experimental results demonstrate that MUN strengthens the reliability of world models and significantly improves the policy's capacity to generalize across new goal settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Context and Diversity Matter: The Emergence of In-Context Learning in World ModelsFan Wang, ZHIYUAN CHEN, YUXUAN ZHONG, Sunjian Zheng 等ICLR 2026 · 被引用 5 次
- Learning from Demonstrations via Capability-Aware Goal SamplingYuanlin Duan, Yuning Wang, Wenjie Qiu, He ZhuNeurIPS 2025 · 被引用 1 次
- Rejecting Hallucinated State Targets during PlanningHarry Zhao, Tristan Sylvain, Romain Laroche, Doina Precup 等ICML 2025
它引用的顶会 Paper24
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel 等ICML 2020 · 被引用 489 次
- Temporal Difference Learning for Model Predictive ControlNicklas Hansen, Hao Su, Xiaolong WangICML 2022 · 被引用 388 次
相关 Paper
- Go Beyond Imagination: Maximizing Episodic Reachability with World ModelsYao Fu, Run Peng, Honglak LeeICML 2023 · 被引用 1 次
- First-Order Representation Languages for Goal-Conditioned RLSimon Ståhlberg, Hector GeffnerAAAI 2026 · 被引用 1 次
- Planning from Pixels using Inverse Dynamics ModelsKeiran Paster, Sheila A. McIlraith, Jimmy BaICLR 2021 · 被引用 44 次
- PlanGAN: Model-based Planning With Sparse Rewards and Multiple GoalsHenry Charlesworth, Giovanni MontanaNeurIPS 2020 · 被引用 34 次
- Active Hierarchical Exploration with Stable Subgoal Representation LearningSiyuan Li, Jin Zhang, Jianhao Wang, Yang Yu 等ICLR 2022 · 被引用 28 次
