Lune

ICML2021Top-tier venue

Safe Reinforcement Learning with Linear Function Approximation

Sanae Amani, Christos Thrampoulidis, Lin Yang

2021Year
42Citations
18Top-tier citations

Abstract

Safety in reinforcement learning has become increasingly important in recent years. Yet, existing solutions either fail to strictly avoid choosing unsafe actions, which may lead to catastrophic results in safety-critical systems, or fail to provide regret guarantees for settings where safety constraints need to be learned. In this paper, we address both problems by first modeling safety as an unknown linear cost function of states and actions, which must always fall below a certain threshold. We then present algorithms, termed SLUCB-QVI and RSLUCB-QVI, for episodic Markov decision processes (MDPs) with linear function approximation. We show that SLUCB-QVI and RSLUCB-QVI, while with no safety violation, achieve a O~(κd3H3T)\tilde{\mathcal{O}}\left(\kappa\sqrt{d^3H^3T}\right) regret, nearly matching that of state-of-the-art unsafe algorithms, where HH is the duration of each episode, dd is the dimension of the feature mapping, κ\kappa is a constant characterizing the safety constraints, and TT is the total number of action plays. We further present numerical simulations that corroborate our theoretical findings.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 0e3c3eab-666f-4162-be47-7765f2c18f2b

Cited by top-tier papers18

Ask how each one uses it

Builds on5

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines