Modeling Unseen Environments with Language-guided Composable Causal Components in Reinforcement Learning
Xinyue Wang, Biwei Huang
Abstract
Generalization in reinforcement learning (RL) remains a significant challenge, especially when agents encounter novel environments with unseen dynamics. Drawing inspiration from human compositional reasoning-where known components are reconfigured to handle new situations-we introduce World Modeling with Compositional Causal Components (WM3C). This novel framework enhances RL generalization by learning and leveraging compositional causal components. Unlike previous approaches focusing on invariant representation learning or metalearning, WM3C identifies and utilizes causal dynamics among composable elements, facilitating robust adaptation to new tasks. Our approach integrates language as a compositional modality to decompose the latent space into meaningful components and provides theoretical guarantees for their unique identification under mild assumptions. Our practical implementation uses a masked autoencoder with mutual information constraints and adaptive sparsity regularization to capture high-level semantic information and effectively disentangle transition dynamics. Experiments on numerical simulations and real-world robotic manipulation tasks demonstrate that WM3C significantly outperforms existing methods in identifying latent processes, improving policy learning, and generalizing to unseen tasks. 1 Proposition 1 (Language-Controlled Components). Under the assumption that the graphical representation of the environment model is Markov and faithful (Spirtes et al., 2001; Glymour et al., 2019; Pearl, 2009) to the data, c i,t 2 is a minimal subset of state dimensions that are directly controlled by the language component l i and s j,t ∈ c i,t if and only if s j,t ̸ ⊥ ⊥ l i | a t-1:t , s t-1 , and s j,t ⊥ ⊥ l k k̸ =i | l i , a t-1:t , s t-1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Ada-Diffuser: Latent-Aware Adaptive Diffusion for Decision-MakingFan Feng, Selena Ge, Minghao Fu, Zijian Li et al.ICLR 2026 · 3 citations
- Plan in Sandbox, Navigate in Open Worlds: Learning Physics-Grounded Abstracted Experience for Embodied NavigationZhixuan Shen, Jiawei Du, Ziyu Guo, Han Luo et al.ICML 2026
- Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured ModelingFan Feng, Yujia Zheng, Minghao Fu, Yongqiang Chen et al.ICML 2026
- World Models in Pieces: Structural Certification for General AgentsYikai Lu, Yifei Wu, Xinyu Lu, Tongxin LiICML 2026
Builds on18
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Self-Supervised Learning with Data Augmentations Provably Isolates Content from StyleJulius von Kügelgen, Yash Sharma, Luigi Gresele, Wieland Brendel et al.NeurIPS 2021 · 421 citations
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 388 citations
- Network Randomization: A Simple Technique for Generalization in Deep Reinforcement LearningKimin Lee, Kibok Lee, Jinwoo Shin, Honglak LeeICLR 2020 · 191 citations
- Towards Nonlinear Disentanglement in Natural Data with Temporal Sparse CodingDavid A. Klindt, Lukas Schott, Yash Sharma, Ivan Ustyuzhaninov et al.ICLR 2021 · 156 citations
Related papers
- MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-ScenariosXuantang Xiong, Ni Mu, Runpeng Xie, Senhao Yang et al.AAAI 2026
- CR³: Boosting Compositional Reasoning in MLLMs Through Rule-Based Reinforcement LearningShun Qian, Bingquan Liu, Chengjie Sun, Peijin Xie et al.AAAI 2026
- Toward Compositional Generalization in Object-Oriented World ModelingLinfeng Zhao, Lingzhi Kong, Robin Walters, Lawson L. S. WongICML 2022 · 28 citations
- Tell me why! Explanations support learning relational and causal structureAndrew K. Lampinen, Nicholas A. Roy, Ishita Dasgupta, Stephanie C. Y. Chan et al.ICML 2022 · 51 citations
- From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model ReasoningLingjing Kong, Xin Liu, Guangyi Chen, Martin Q. Ma et al.ICML 2026 · 1 citation
