Curious Exploration via Structured World Models Yields Zero-Shot Object Manipulation
Cansu Sancaktar, Sebastian Blaes, Georg Martius
摘要
It has been a long-standing dream to design artificial agents that explore their environment efficiently via intrinsic motivation, similar to how children perform curious free play. Despite recent advances in intrinsically motivated reinforcement learning (RL), sample-efficient exploration in object manipulation scenarios remains a significant challenge as most of the relevant information lies in the sparse agent-object and object-object interactions. In this paper, we propose to use structured world models to incorporate relational inductive biases in the control loop to achieve sample-efficient and interaction-rich exploration in compositional multi-object environments. By planning for future novelty inside structured world models, our method generates free-play behavior that starts to interact with objects early on and develops more complex behavior over time. Instead of using models only to compute intrinsic rewards, as commonly done, our method showcases that the self-reinforcing cycle between good models and good exploration also opens up another avenue: zero-shot generalization to downstream tasks via model-based planning. After the entirely intrinsic task-agnostic exploration phase, our method solves challenging downstream tasks such as stacking, flipping, pick & place, and throwing and generalizes to unseen numbers and arrangements of objects without any additional training. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Optimistic Active Exploration of Dynamical SystemsBhavya Sukhija, Lenart Treven, Cansu Sancaktar, Sebastian Blaes 等NeurIPS 2023 · 被引用 42 次
- Entity-Centric Reinforcement Learning for Object Manipulation from PixelsDan Haramati, Tal Daniel, Aviv TamarICLR 2024 · 被引用 31 次
- Learning Hierarchical World Models with Adaptive Temporal Abstractions from Discrete Latent DynamicsChristian Gumbsch, Noor Sajid, Georg Martius, Martin V. ButzICLR 2024 · 被引用 24 次
- Disentangled Unsupervised Skill Discovery for Efficient Hierarchical Reinforcement LearningJiaheng Hu, Zizhao Wang, Peter Stone, Roberto Martín-MartínNeurIPS 2024 · 被引用 21 次
- Emergent Dexterity Via Diverse Resets and Large-Scale Reinforcement LearningPatrick Yin, Tyler Westenbroek, Zhengyu Zhang, Ignacio Dagnino 等ICLR 2026 · 被引用 15 次
它引用的顶会 Paper8
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Learning to Simulate Complex Physics with Graph NetworksAlvaro Sanchez-Gonzalez, Jonathan Godwin, Tobias Pfaff, Rex Ying 等ICML 2020 · 被引用 1,439 次
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel 等ICML 2020 · 被引用 489 次
- Contrastive Learning of Structured World ModelsThomas N. Kipf, Elise van der Pol, Max WellingICLR 2020 · 被引用 322 次
- Discovering and Achieving Goals via World ModelsRussell Mendonca, Oleh Rybkin, Kostas Daniilidis, Danijar Hafner 等NeurIPS 2021 · 被引用 177 次
相关 Paper
- Regularity as Intrinsic Reward for Free PlayCansu Sancaktar, Justus H. Piater, Georg MartiusNeurIPS 2023 · 被引用 9 次
- Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured ModelingFan Feng, Yujia Zheng, Minghao Fu, Yongqiang Chen 等ICML 2026
- Novelty Search in Representational Space for Sample Efficient ExplorationRuo Yu Tao, Vincent François-Lavet, Joelle PineauNeurIPS 2020 · 被引用 53 次
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated EnvironmentsRoberta Raileanu, Tim RocktäschelICLR 2020 · 被引用 198 次
- SENSEI: Semantic Exploration Guided by Foundation Models to Learn Versatile World ModelsCansu Sancaktar, Christian Gumbsch, Andrii Zadaianchuk, Pavel Kolev 等ICML 2025
