Lune

ICML2023顶会

Robust Situational Reinforcement Learning in Face of Context Disturbances

Jinpeng Zhang, Yufeng Zheng, Chuheng Zhang, Li Zhao, Lei Song, Yuan Zhou, Jiang Bian

出版方
2023年份
5被引次数
2顶会引用

摘要

In many real-world tasks, the presence of dynamic and uncontrollable environmental factors, commonly referred to as context, plays a crucial role in the decision-making process. Examples of such factors include customer demand in inventory control and the speed of the lead car in autonomous driving. One of the challenges of reinforcement learning in these applications is that the true context transitions can be easily exposed to some unknown source of contamination, leading to a shift of context transitions between source domains and target domains, which could cause performance degradation for RL algorithms. To tackle this problem, we propose the robust situational Markov decision process (RS-MDP) framework which captures the possible deviations of context transitions explicitly. To scale to large context space, we introduce the softmin smoothed robust Bellman operator to learn the robust Q-value approximately, and extend existing RL algorithm SAC to learn the desired robust policies under our RS-MDP framework. We conduct experiments on several locomotion tasks with dynamic contexts and inventory control tasks to demonstrate that our algorithm can generalize better and be more robust against context disturbances, and outperform existing basic RL algorithms that do not consider robustness and robust RL algorithms that consider robustness over the whole state transitions.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

它引用的顶会 Paper10

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖