Modeling Unseen Environments with Language-guided Composable Causal Components in Reinforcement Learning
Xinyue Wang, Biwei Huang
摘要
Generalization in reinforcement learning (RL) remains a significant challenge, especially when agents encounter novel environments with unseen dynamics. Drawing inspiration from human compositional reasoning-where known components are reconfigured to handle new situations-we introduce World Modeling with Compositional Causal Components (WM3C). This novel framework enhances RL generalization by learning and leveraging compositional causal components. Unlike previous approaches focusing on invariant representation learning or metalearning, WM3C identifies and utilizes causal dynamics among composable elements, facilitating robust adaptation to new tasks. Our approach integrates language as a compositional modality to decompose the latent space into meaningful components and provides theoretical guarantees for their unique identification under mild assumptions. Our practical implementation uses a masked autoencoder with mutual information constraints and adaptive sparsity regularization to capture high-level semantic information and effectively disentangle transition dynamics. Experiments on numerical simulations and real-world robotic manipulation tasks demonstrate that WM3C significantly outperforms existing methods in identifying latent processes, improving policy learning, and generalizing to unseen tasks. 1 Proposition 1 (Language-Controlled Components). Under the assumption that the graphical representation of the environment model is Markov and faithful (Spirtes et al., 2001; Glymour et al., 2019; Pearl, 2009) to the data, c i,t 2 is a minimal subset of state dimensions that are directly controlled by the language component l i and s j,t ∈ c i,t if and only if s j,t ̸ ⊥ ⊥ l i | a t-1:t , s t-1 , and s j,t ⊥ ⊥ l k k̸ =i | l i , a t-1:t , s t-1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Ada-Diffuser: Latent-Aware Adaptive Diffusion for Decision-MakingFan Feng, Selena Ge, Minghao Fu, Zijian Li 等ICLR 2026 · 被引用 3 次
- Plan in Sandbox, Navigate in Open Worlds: Learning Physics-Grounded Abstracted Experience for Embodied NavigationZhixuan Shen, Jiawei Du, Ziyu Guo, Han Luo 等ICML 2026
- Learning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured ModelingFan Feng, Yujia Zheng, Minghao Fu, Yongqiang Chen 等ICML 2026
- World Models in Pieces: Structural Certification for General AgentsYikai Lu, Yifei Wu, Xinyu Lu, Tongxin LiICML 2026
它引用的顶会 Paper18
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Self-Supervised Learning with Data Augmentations Provably Isolates Content from StyleJulius von Kügelgen, Yash Sharma, Luigi Gresele, Wieland Brendel 等NeurIPS 2021 · 被引用 421 次
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 被引用 388 次
- Network Randomization: A Simple Technique for Generalization in Deep Reinforcement LearningKimin Lee, Kibok Lee, Jinwoo Shin, Honglak LeeICLR 2020 · 被引用 191 次
- Towards Nonlinear Disentanglement in Natural Data with Temporal Sparse CodingDavid A. Klindt, Lukas Schott, Yash Sharma, Ivan Ustyuzhaninov 等ICLR 2021 · 被引用 156 次
相关 Paper
- MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-ScenariosXuantang Xiong, Ni Mu, Runpeng Xie, Senhao Yang 等AAAI 2026
- CR³: Boosting Compositional Reasoning in MLLMs Through Rule-Based Reinforcement LearningShun Qian, Bingquan Liu, Chengjie Sun, Peijin Xie 等AAAI 2026
- Toward Compositional Generalization in Object-Oriented World ModelingLinfeng Zhao, Lingzhi Kong, Robin Walters, Lawson L. S. WongICML 2022 · 被引用 28 次
- Tell me why! Explanations support learning relational and causal structureAndrew K. Lampinen, Nicholas A. Roy, Ishita Dasgupta, Stephanie C. Y. Chan 等ICML 2022 · 被引用 51 次
- From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model ReasoningLingjing Kong, Xin Liu, Guangyi Chen, Martin Q. Ma 等ICML 2026 · 被引用 1 次
