Tackling Non-Stationarity in Reinforcement Learning via Causal-Origin Representation
Wanpeng Zhang, Yilin Li, Boyu Yang, Zongqing Lu
Abstract
In real-world scenarios, the application of reinforcement learning is significantly challenged by complex non-stationarity. Most existing methods attempt to model changes in the environment explicitly, often requiring impractical prior knowledge of environments. In this paper, we propose a new perspective, positing that non-stationarity can propagate and accumulate through complex causal relationships during state transitions, thereby compounding its sophistication and affecting policy learning. We believe that this challenge can be more effectively addressed by implicitly tracing the causal origin of non-stationarity. To this end, we introduce the Causal-Origin REPresentation (COREP) algorithm. COREP primarily employs a guided updating mechanism to learn a stable graph representation for the state, termed as causal-origin representation. By leveraging this representation, the learned policy exhibits impressive resilience to non-stationarity. We supplement our approach with a theoretical analysis grounded in the causal interpretation for non-stationary reinforcement learning, advocating for the validity of the causal-origin representation. Experimental results further demonstrate the superior performance of COREP over existing methods in tackling non-stationarity problems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9ccc9af-52dc-4cd2-857a-98bd1924c0d9Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Domain Adaptation as a Problem of Inference on Graphical ModelsKun Zhang, Mingming Gong, Petar Stojanov, Biwei Huang et al.NeurIPS 2020 · 76 citations
- AdaRL: What, Where, and How to Adapt in Transfer Reinforcement LearningBiwei Huang, Fan Feng, Chaochao Lu, Sara Magliacane et al.ICLR 2022 · 75 citations
- Optimizing for the Future in Non-Stationary MDPsYash Chandak, Georgios Theocharous, Shiv Shankar, Martha White et al.ICML 2020 · 72 citations
- Factored Adaptation for Non-Stationary Reinforcement LearningFan Feng, Biwei Huang, Kun Zhang, Sara MagliacaneNeurIPS 2022 · 52 citations
- Analyzing the Expressive Power of Graph Neural Networks in a Spectral PerspectiveMuhammet Balcilar, Guillaume Renton, Pierre Héroux, Benoit Gaüzère et al.ICLR 2021 · 44 citations
Related papers
- Towards Generalizable Reinforcement Learning via Causality-Guided Self-Adaptive RepresentationsYupei Yang, Biwei Huang, Fan Feng, Xinyue Wang et al.ICLR 2025
- Deep Reinforcement Learning amidst Continual Structured Non-StationarityAnnie Xie, James Harrison, Chelsea FinnICML 2021 · 43 citations
- COGS: A Causal Representation Learning Framework for Out-of-Distribution Generalization in Time SeriesXinxin Song, Yuxiao Cheng, Tingxiong Xiao, Jinli SuoAAAI 2026
- Causal Temporal Representation Learning with Nonstationary Sparse TransitionXiangchen Song, Zijian Li, Guangyi Chen, Yujia Zheng et al.NeurIPS 2024 · 18 citations
- Curious Causality-Seeking Agents in Open-ended WorldsZhiyu Zhao, Haoxuan Li, Haifeng Zhang, Jun Wang et al.NeurIPS 2025
