Lune

NeurIPS2022Top-tier venue

Provably Feedback-Efficient Reinforcement Learning via Active Reward Learning

Dingwen Kong, Lin Yang

2022Year
19Citations
4Top-tier citations

Abstract

An appropriate reward function is of paramount importance in specifying a task in reinforcement learning (RL). Yet, it is known to be extremely challenging in practice to design a correct reward function for even simple tasks. Human-in-the-loop (HiL) RL allows humans to communicate complex goals to the RL agent by providing various types of feedback. However, despite achieving great empirical successes, HiL RL usually requires too much feedback from a human teacher and also suffers from insufficient theoretical understanding. In this paper, we focus on addressing this issue from a theoretical perspective, aiming to provide provably feedback-efficient algorithmic frameworks that take human-in-the-loop to specify rewards of given tasks. We provide an active-learning-based RL algorithm that first explores the environment without specifying a reward function and then asks a human teacher for only a few queries about the rewards of a task at some state-action pairs. After that, the algorithm guarantees to provide a nearly optimal policy for the task with high probability. We show that, even with the presence of random noise in the feedback, the algorithm only takes O~(Hdim⁡R2)\widetilde{O}(H{{\dim_{R}^2}}) queries on the reward function to provide an ϵ\epsilon-optimal policy for any ϵ>0\epsilon>0. Here HH is the horizon of the RL environment, and dim⁡R\dim_{R} specifies the complexity of the function class representing the reward function. In contrast, standard RL algorithms require to query the reward function for at least Ω(poly⁡(d,1/ϵ))\Omega(\operatorname{poly}(d, 1/\epsilon)) state-action pairs where dd depends on the complexity of the environmental transition.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 188d5c05-b21c-4d32-82c8-593a9ceffcb6

Cited by top-tier papers4

Ask how each one uses it

Builds on25

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines