Lune

NeurIPS2025Top-tier venue

EDELINE: Enhancing Memory in Diffusion-based World Models via Linear-Time Sequence Modeling

Jia-Hua Lee, Bor-Jiun Lin, Wei-Fang Sun, Chun-Yi Lee

2025Year
4Citations

Abstract

World models represent a promising approach for training reinforcement learning agents with significantly improved sample efficiency. While most world model methods primarily rely on sequences of discrete latent variables to model environment dynamics, this compression often neglects critical visual details essential for reinforcement learning. Recent diffusion-based world models condition generation on a fixed context length of frames to predict the next observation, using separate recurrent neural networks to model rewards and termination signals. Although this architecture effectively enhances visual fidelity, the fixed context length approach inherently limits memory capacity. In this paper, we introduce EDELINE, a unified world model architecture that integrates state space models with diffusion models. Our approach outperforms existing baselines across visually challenging Atari 100k tasks, memory-demanding Crafter benchmark, and 3D first-person ViZDoom environments, demonstrating superior performance in all these diverse challenges. Code is available at https://github.com/LJH-coding/EDELINE.

recent work such as R2I [14] has addressed fundamental challenges in long-term memory and credit assignment, demonstrating the importance of enhanced temporal reasoning capabilities.

Based on these considerations, we introduce EDELINE (Enhancing Diffusion-basEd World Models via LINEar-Time Sequence Modeling), a unified framework that integrates the advantages of diffusion models and SSMs. EDELINE advances the state-of-the-art (SOTA) through three key innovations:

(1) Memory Enhancement: A recurrent embedding module (REM) based on Mamba SSMs that processes unbounded observation-action sequences to enable adaptive memory retention beyond fixed-context limitations, (2) Unified Framework: Direct conditioning of reward and termination prediction on REM hidden states that eliminates separate networks for efficient representation sharing, and (3) Dynamic Loss Harmonization: Adaptive weighting of observation and reward losses that addresses scale disparities in multi-task optimization. To validate EDELINE's effectiveness, we conduct extensive evaluations. It achieves 1.87× human normalized scores on the sample-efficiency challenging Atari 100k [15] benchmark and surpasses all model-based methods that do not use look-ahead search. The ablation studies confirm the effectiveness of each architectural component for world modeling performance. To substantiate EDELINE's capacity to preserve long temporal information for consistent predictions, we evaluate performance on environments that require longterm memory capability: MiniGrid-Memory [16], Crafter [17], and ViZDoom [18]. Both qualitative and quantitative results show superior temporal consistency in modeling and imaginary quality across 2D and 3D environments. Furthermore, EDELINE shows significant performance improvements compared to prior diffusion-based world models. Our contributions can be summarized as follows:

• We introduce a unified architecture EDELINE that integrates a Next-Frame Predictor for future observation imaginary, a Recurrent Embedding Module for temporal sequence processing, and a Reward/Termination Predictor to address the long-term memory limitations in existing models.

• EDELINE utilizes an SSMs-based embedding module that overcomes the fixed context limitations of prior diffusion-based methods and enhances performance in memory-demanding environments.

• Our experimental insights across various benchmarks validate EDELINE's precise prediction of reward-critical elements where prior diffusion-based world models exhibit structural inaccuracies.

2 Related Work

Diffusion models have revolutionized high-resolution image generation through their noise-reversal process. Foundational works including DDPM [19] and DDIM [20] established core principles for subsequent developments. Score-based models [21,22] enhanced sampling efficiency through gradient estimation of data distributions, while energy-based models [23] introduced robust optimization properties via probabilistic state modeling. The application of diffusion models in RL has expanded significantly. These models serve as policy networks for efficient offline learning [24,25,26], enable diverse strategy generation in planning tasks [27,28], and provide novel approaches to reward modeling [29]. MetaDiffuser [30] showed effectiveness as conditional planners in offline meta-RL, while others have adopted diffusion models for trajectory modeling and synthetic experience generation [31].

World models serve as a fundamental component in model-based RL, enabling sample-efficient and safe learning through simulated environments. SimPLe [15] established the groundwork by introducing world models to the Atari domain and proposing the Atari 100k benchmark. Dreamer [6] advanced this field through RL from latent space predictions, which DreamerV2 [7] further refined with discrete latents to mitigate comp

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Builds on47

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines