Lune

NeurIPS2021Top-tier venue

Explicable Reward Design for Reinforcement Learning Agents

Rati Devidze, Goran Radanovic, Parameswaran Kamalaruban, Adish Singla

2021Year
60Citations
20Top-tier citations

Abstract

We study the design of explicable reward functions for a reinforcement learning agent while guaranteeing that an optimal policy induced by the function belongs to a set of target policies. By being explicable, we seek to capture two properties: (a) informativeness so that the rewards speed up the agent's convergence, and (b) sparseness as a proxy for ease of interpretability of the rewards. The key challenge is that higher informativeness typically requires dense rewards for many learning tasks, and existing techniques do not allow one to balance these two properties appropriately. In this paper, we investigate the problem from the perspective of discrete optimization and introduce a novel framework, EXPRD, to design explicable reward functions. EXPRD builds upon an informativeness criterion that captures the (sub-)optimality of target policies at different time horizons in terms of actions taken from any given starting state. We provide a mathematical analysis of EXPRD, and show its connections to existing reward design techniques, including potential-based reward shaping. Experimental results on two navigation tasks demonstrate the effectiveness of EXPRD in designing explicable reward functions. In addition, for any state s ∈ S, globally optimal actions Π * s ⊆ A under R are also myopically optimal under R PBRS since δ π 0 (s, a) = δ * ∞ (s, a) for all π ∈ Π * [3, 8] -this leads to a dramatic speed-up in the learning process. However, the potentialbased reward shaping produces dense reward function which is less interpretable (see Section 4). * and the ∞-step optimality gaps by δ * ∞ ; the quantities defined corresponding to R := R are denoted by a widehat, e.g., the optimal policy set by Π * .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext e967ec07-c433-4bb9-b825-58707e3f8461

Cited by top-tier papers20

Ask how each one uses it

Builds on5

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines