Thinker: Learning to Plan and Act
Stephen Chung, Ivan Anokhin, David Krueger
摘要
We propose the Thinker algorithm, a novel approach that enables reinforcement learning agents to autonomously interact with and utilize a learned world model. The Thinker algorithm wraps the environment with a world model and introduces new actions designed for interacting with the world model. These model-interaction actions enable agents to perform planning by proposing alternative plans to the world model before selecting a final action to execute in the environment. This approach eliminates the need for handcrafted planning algorithms by enabling the agent to learn how to plan autonomously and allows for easy interpretation of the agent's plan with visualization. We demonstrate the algorithm's effectiveness through experimental results in the game of Sokoban and the Atari 2600 benchmark, where the Thinker algorithm achieves state-of-the-art performance and competitive results, respectively. Visualizations of agents trained with the Thinker algorithm demonstrate that they have learned to plan effectively with the world model to select better actions. Thinker is the first work showing that an RL agent can learn to plan with a learned world model in complex environments. 1 Full code is available at https://github.com/stephen-chung-mh/thinker , which allows for using the Thinker-augmented MDP with the same interface as OpenAI Gym [10] . 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the EnvironmentHao Tang, Darren Key, Kevin EllisNeurIPS 2024 · 被引用 123 次
- Reinforcement Learning with Lookahead InformationNadav MerlisNeurIPS 2024 · 被引用 11 次
- Thinker: Learning to Think Fast and SlowStephen Chung, Wenyu Du, Jie FuNeurIPS 2025 · 被引用 10 次
- When Can Model-Free Reinforcement Learning be Enough for Thinking?Josiah Hanna, Nicholas CorradoNeurIPS 2025 · 被引用 6 次
- Deep RL Needs Deep Behavior Analysis: Exploring Implicit Planning by Model-Free Agents in Open-Ended EnvironmentsRiley Simmons-Edler, Ryan Paul Badman, Felix Baastad Berg, Raymond Chua 等NeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper10
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- Stabilizing Transformers for Reinforcement LearningEmilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu 等ICML 2020 · 被引用 464 次
- Mastering Atari Games with Limited DataWeirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel 等NeurIPS 2021 · 被引用 345 次
- A Self-Tuning Actor-Critic AlgorithmTom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel 等NeurIPS 2020 · 被引用 106 次
- Meta-Gradient Reinforcement Learning with an Objective Discovered OnlineZhongwen Xu, Hado Philip van Hasselt, Matteo Hessel, Junhyuk Oh 等NeurIPS 2020 · 被引用 90 次
相关 Paper
- Interpreting Emergent Planning in Model-Free Reinforcement LearningThomas Bush, Stephen Chung, Usman Anwar, Adrià Garriga-Alonso 等ICLR 2025
- Thinker: Training LLMs in Hierarchical Thinking for Deep Search via Multi-Turn InteractionJun Xu, Xinkai Du, Yu Ao, Peilong Zhao 等AAAI 2026 · 被引用 3 次
- Learning World Models for Unconstrained Goal NavigationYuanlin Duan, Wensen Mao, He ZhuNeurIPS 2024 · 被引用 11 次
- Reader: Model-based language-instructed reinforcement learningNicola Dainese, Pekka Marttinen, Alexander IlinEMNLP 2023 · 被引用 1 次
- Learning Transformer-based World Models with Contrastive Predictive CodingMaxime Burchi, Radu TimofteICLR 2025
