Lune

NeurIPS2022Top-tier venue

Exploration-Guided Reward Shaping for Reinforcement Learning under Sparse Rewards

Rati Devidze, Parameswaran Kamalaruban, Adish Singla

2022Year
122Citations
21Top-tier citations

Abstract

We study the problem of reward shaping to accelerate the training process of a reinforcement learning agent. Existing works have considered a number of different reward shaping formulations; however, they either require external domain knowledge or fail in environments with extremely sparse rewards. In this paper, we propose a novel framework, Exploration-Guided Reward Shaping (E XPLO RS), that operates in a fully self-supervised manner and can accelerate an agent’s learning even in sparse-reward environments. The key idea of E XPLO RS is to learn an intrinsic reward function in combination with exploration-based bonuses to maximize the agent’s utility w.r.t. extrinsic rewards. We theoretically showcase the usefulness of our reward shaping framework in a special family of MDPs. Experimental results on several environments with sparse/noisy reward signals demonstrate the effectiveness of E XPLO RS.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 7e37204f-90f7-4854-a93d-d234658f109f

Cited by top-tier papers21

Ask how each one uses it

Builds on3

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines