Identifying Latent State-Transition Processes for Individualized Reinforcement Learning
Yuewen Sun, Biwei Huang, Yu Yao, Donghuo Zeng, Xinshuai Dong, Songyao Jin, Boyang Sun, Roberto Legaspi, Kazushi Ikeda, Peter Spirtes, Kun Zhang
Abstract
The application of reinforcement learning (RL) involving interactions with individuals has grown significantly in recent years. These interactions, influenced by factors such as personal preferences and physiological differences, causally influence state transitions, ranging from health conditions in healthcare to learning progress in education. As a result, different individuals may exhibit different state-transition processes. Understanding individualized state-transition processes is essential for optimizing individualized policies. In practice, however, identifying these state-transition processes is challenging, as individual-specific factors often remain latent. In this paper, we establish the identifiability of these latent factors and introduce a practical method that effectively learns these processes from observed state-action trajectories. Experiments on various datasets show that the proposed method can effectively identify latent state-transition processes and facilitate the learning of individualized RL policies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b2bc8cc-935d-42a2-b0a6-682e32117741Cited by top-tier papers2
- Efficient and Trustworthy Causal Discovery with Latent Variables and Complex RelationsXiu-Chuan Li, Tongliang LiuICLR 2025
- Recovery of Causal Graph Involving Latent Variables via Homologous SurrogatesXiu-Chuan Li, Jun Wang, Tongliang LiuICLR 2025
Builds on23
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski et al.ICLR 2020 · 969 citations
- Personalized HeartSteps: A Reinforcement Learning Algorithm for Optimizing Physical ActivityPeng Liao, Kristjan H. Greenewald, Predrag V. Klasnja, Susan A. MurphyUbiComp 2020 · 163 citations
- Invariant Causal Prediction for Block MDPsAmy Zhang, Clare Lyle, Shagun Sodhani, Angelos Filos et al.ICML 2020 · 153 citations
- Bootstrap Latent-Predictive Representations for Multitask Reinforcement LearningZhaohan Daniel Guo, Bernardo Ávila Pires, Bilal Piot, Jean-Bastien Grill et al.ICML 2020 · 153 citations
Related papers
- Temporally Disentangled Representation LearningWeiran Yao, Guangyi Chen, Kun ZhangNeurIPS 2022 · 84 citations
- Learning World Models with Identifiable FactorizationYuren Liu, Biwei Huang, Zhengmao Zhu, Hong-Long Tian et al.NeurIPS 2023 · 32 citations
- Factored Adaptation for Non-Stationary Reinforcement LearningFan Feng, Biwei Huang, Kun Zhang, Sara MagliacaneNeurIPS 2022 · 52 citations
- Latent Learning Progress Drives Autonomous Goal Selection in Human Reinforcement LearningGaia Molinaro, Cédric Colas, Pierre-Yves Oudeyer, Anne CollinsNeurIPS 2024 · 12 citations
- Towards Safe Policy Learning under Partial Identifiability: A Causal ApproachShalmali Joshi, Junzhe Zhang, Elias BareinboimAAAI 2024 · 10 citations
