Identifiable Token Correspondence for World Models
Youngin Kim, Ray Sun, Inho Kim, Bumsoo Park, Hyun Oh Song
Abstract
Token-based transformer world models have shown strong performance in visual reinforcement learning, but often suffer from temporal inconsistency in long-horizon rollouts, including object duplication, disappearance, and transmutation. A key reason is that most existing approaches treat next-frame prediction purely as a token generation problem, without considering the persistence of tokens across time. We introduce Identifiable Token Correspondence (ITC), a decoding step for token-based transformer world models that formulates next-frame prediction as a structured assignment problem with latent token correspondence variables: each next-frame token is explained either by copying a token from the previous frame or by generating a new one. ITC leaves the transformer architecture and training procedure unchanged and can be added on top of existing backbones. Our experiments show state-of-the-art performance on 4 challenging benchmarks. The proposed method achieves a return of 72.5% and a score of 35.6% on the Craftax-classic benchmark, significantly surpassing the previous best of 67.4% and 27.9%. We release our source code on https://github.com/snu-mllab/ Identifiable-Token-Correspondence .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88339bb1-a161-469b-a9f4-0b86864d3608Builds on18
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
- Puzzle Mix: Exploiting Saliency and Local Statistics for Optimal MixupJang-Hyun Kim, Wonho Choo, Hyun Oh SongICML 2020 · 457 citations
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 388 citations
- Diffusion for World Modeling: Visual Details Matter in AtariEloi Alonso, Adam Jelley, Vincent Micheli, Anssi Kanervisto et al.NeurIPS 2024 · 359 citations
Related papers
- Improving Transformer World Models for Data-Efficient RLAntoine Dedieu, Joseph Ortiz, Xinghua Lou, Carter Wendelken et al.ICML 2025
- Learning Transformer-based World Models with Contrastive Predictive CodingMaxime Burchi, Radu TimofteICLR 2025
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
- Efficient World Models with Context-Aware TokenizationVincent Micheli, Eloi Alonso, François FleuretICML 2024 · 25 citations
- When Do Transformers Shine in RL? Decoupling Memory from Credit AssignmentTianwei Ni, Michel Ma, Benjamin Eysenbach, Pierre-Luc BaconNeurIPS 2023 · 77 citations
