Identifying Latent State-Transition Processes for Individualized Reinforcement Learning
Yuewen Sun, Biwei Huang, Yu Yao, Donghuo Zeng, Xinshuai Dong, Songyao Jin, Boyang Sun, Roberto Legaspi, Kazushi Ikeda, Peter Spirtes, Kun Zhang
摘要
The application of reinforcement learning (RL) involving interactions with individuals has grown significantly in recent years. These interactions, influenced by factors such as personal preferences and physiological differences, causally influence state transitions, ranging from health conditions in healthcare to learning progress in education. As a result, different individuals may exhibit different state-transition processes. Understanding individualized state-transition processes is essential for optimizing individualized policies. In practice, however, identifying these state-transition processes is challenging, as individual-specific factors often remain latent. In this paper, we establish the identifiability of these latent factors and introduce a practical method that effectively learns these processes from observed state-action trajectories. Experiments on various datasets show that the proposed method can effectively identify latent state-transition processes and facilitate the learning of individualized RL policies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Efficient and Trustworthy Causal Discovery with Latent Variables and Complex RelationsXiu-Chuan Li, Tongliang LiuICLR 2025
- Recovery of Causal Graph Involving Latent Variables via Homologous SurrogatesXiu-Chuan Li, Jun Wang, Tongliang LiuICLR 2025
它引用的顶会 Paper23
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 被引用 1,170 次
- Model Based Reinforcement Learning for AtariLukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski 等ICLR 2020 · 被引用 969 次
- Personalized HeartSteps: A Reinforcement Learning Algorithm for Optimizing Physical ActivityPeng Liao, Kristjan H. Greenewald, Predrag V. Klasnja, Susan A. MurphyUbiComp 2020 · 被引用 163 次
- Invariant Causal Prediction for Block MDPsAmy Zhang, Clare Lyle, Shagun Sodhani, Angelos Filos 等ICML 2020 · 被引用 153 次
- Bootstrap Latent-Predictive Representations for Multitask Reinforcement LearningZhaohan Daniel Guo, Bernardo Ávila Pires, Bilal Piot, Jean-Bastien Grill 等ICML 2020 · 被引用 153 次
相关 Paper
- Temporally Disentangled Representation LearningWeiran Yao, Guangyi Chen, Kun ZhangNeurIPS 2022 · 被引用 84 次
- Learning World Models with Identifiable FactorizationYuren Liu, Biwei Huang, Zhengmao Zhu, Hong-Long Tian 等NeurIPS 2023 · 被引用 32 次
- Factored Adaptation for Non-Stationary Reinforcement LearningFan Feng, Biwei Huang, Kun Zhang, Sara MagliacaneNeurIPS 2022 · 被引用 52 次
- Latent Learning Progress Drives Autonomous Goal Selection in Human Reinforcement LearningGaia Molinaro, Cédric Colas, Pierre-Yves Oudeyer, Anne CollinsNeurIPS 2024 · 被引用 12 次
- Towards Safe Policy Learning under Partial Identifiability: A Causal ApproachShalmali Joshi, Junzhe Zhang, Elias BareinboimAAAI 2024 · 被引用 10 次
