Do Transformer World Models Give Better Policy Gradients?
Michel Ma, Tianwei Ni, Clement Gehring, Pierluca D'Oro, Pierre-Luc Bacon
Abstract
A natural approach for reinforcement learning is to predict future rewards by unrolling a neural network world model, and to backpropagate through the resulting computational graph to learn a policy. However, this method often becomes impractical for long horizons since typical world models induce hard-to-optimize loss landscapes. Transformers are known to efficiently propagate gradients over long horizons: could they be the solution to this problem? Surprisingly, we show that commonly-used transformer world models produce circuitous gradient paths, which can be detrimental to long-range policy gradients. To tackle this challenge, we propose a class of world models called Actions World Models (AWMs), designed to provide more direct routes for gradient propagation. We integrate such AWMs into a policy gradient framework that underscores the relationship between network architectures and the policy gradient updates they inherently represent. We demonstrate that AWMs can generate optimization landscapes that are easier to navigate even when compared to those from the simulator itself. This property allows transformer AWMs to produce better policies than competitive baselines in realistic long-horizon tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 891f022e-1149-4d77-b044-84c66e4844ceCited by top-tier papers4
- DMWM: Dual-Mind World Model with Long-Term ImaginationLingyi Wang, Rashed Shelim, Walid Saad, Naren RamakrishnanNeurIPS 2025 · 15 citations
- The Surprising Difficulty of Search in Model-Based Reinforcement LearningWei-Di Chang, Mikael Henaff, Brandon Amos, Gregory Dudek et al.ICML 2026 · 4 citations
- Long-Horizon Model-Based Offline Reinforcement Learning Without Explicit ConservatismTianwei Ni, Esther Derman, Vineet Jain, Vincent Taboga et al.ICML 2026 · 1 citation
- Reparameterization Proximal Policy OptimizationHai Zhong, Xun Wang, Zhuoran Li, Longbo HuangICML 2026
Builds on19
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 950 citations
Related papers
- Transformer-based World Models Are Happy With 100k InteractionsJan Robine, Marc Höftmann, Tobias Uelwer, Stefan HarmelingICLR 2023 · 4 citations
- Facing Off World Model Backbones: RNNs, Transformers, and S4Fei Deng, Junyeong Park, Sungjin AhnNeurIPS 2023 · 53 citations
- Memo: Training Memory-Efficient Embodied Agents with Reinforcement LearningGunshi Gupta, Karmesh Yadav, Zsolt Kira, Yarin Gal et al.NeurIPS 2025 · 9 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Unlocking Efficient Vehicle Dynamics Modeling via Analytic World ModelsAsen Nachkov, Danda Pani Paudel, Jan-Nico Zaech, Davide Scaramuzza et al.AAAI 2026 · 2 citations
