Thinker: Learning to Plan and Act
Stephen Chung, Ivan Anokhin, David Krueger
Abstract
We propose the Thinker algorithm, a novel approach that enables reinforcement learning agents to autonomously interact with and utilize a learned world model. The Thinker algorithm wraps the environment with a world model and introduces new actions designed for interacting with the world model. These model-interaction actions enable agents to perform planning by proposing alternative plans to the world model before selecting a final action to execute in the environment. This approach eliminates the need for handcrafted planning algorithms by enabling the agent to learn how to plan autonomously and allows for easy interpretation of the agent's plan with visualization. We demonstrate the algorithm's effectiveness through experimental results in the game of Sokoban and the Atari 2600 benchmark, where the Thinker algorithm achieves state-of-the-art performance and competitive results, respectively. Visualizations of agents trained with the Thinker algorithm demonstrate that they have learned to plan effectively with the world model to select better actions. Thinker is the first work showing that an RL agent can learn to plan with a learned world model in complex environments. 1 Full code is available at https://github.com/stephen-chung-mh/thinker , which allows for using the Thinker-augmented MDP with the same interface as OpenAI Gym [10] . 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb36328f-0732-4be1-bbaa-0f64e8c60db4Cited by top-tier papers7
- WorldCoder, a Model-Based LLM Agent: Building World Models by Writing Code and Interacting with the EnvironmentHao Tang, Darren Key, Kevin EllisNeurIPS 2024 · 123 citations
- Reinforcement Learning with Lookahead InformationNadav MerlisNeurIPS 2024 · 11 citations
- Thinker: Learning to Think Fast and SlowStephen Chung, Wenyu Du, Jie FuNeurIPS 2025 · 10 citations
- When Can Model-Free Reinforcement Learning be Enough for Thinking?Josiah Hanna, Nicholas CorradoNeurIPS 2025 · 6 citations
- Deep RL Needs Deep Behavior Analysis: Exploring Implicit Planning by Model-Free Agents in Open-Ended EnvironmentsRiley Simmons-Edler, Ryan Paul Badman, Felix Baastad Berg, Raymond Chua et al.NeurIPS 2025 · 6 citations
Builds on10
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Stabilizing Transformers for Reinforcement LearningEmilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu et al.ICML 2020 · 464 citations
- Mastering Atari Games with Limited DataWeirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel et al.NeurIPS 2021 · 345 citations
- A Self-Tuning Actor-Critic AlgorithmTom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel et al.NeurIPS 2020 · 106 citations
- Meta-Gradient Reinforcement Learning with an Objective Discovered OnlineZhongwen Xu, Hado Philip van Hasselt, Matteo Hessel, Junhyuk Oh et al.NeurIPS 2020 · 90 citations
Related papers
- Interpreting Emergent Planning in Model-Free Reinforcement LearningThomas Bush, Stephen Chung, Usman Anwar, Adrià Garriga-Alonso et al.ICLR 2025
- Thinker: Training LLMs in Hierarchical Thinking for Deep Search via Multi-Turn InteractionJun Xu, Xinkai Du, Yu Ao, Peilong Zhao et al.AAAI 2026 · 3 citations
- Learning World Models for Unconstrained Goal NavigationYuanlin Duan, Wensen Mao, He ZhuNeurIPS 2024 · 11 citations
- Reader: Model-based language-instructed reinforcement learningNicola Dainese, Pekka Marttinen, Alexander IlinEMNLP 2023 · 1 citation
- Learning Transformer-based World Models with Contrastive Predictive CodingMaxime Burchi, Radu TimofteICLR 2025
