Lune

ICML2023Top-tier venue

Model-based Offline Reinforcement Learning with Count-based Conservatism

Byeongchan Kim, Min-hwan Oh

2023Year
19Citations
8Top-tier citations

Abstract

In this paper, we propose a model-based offline reinforcement learning method that integrates count-based conservatism, named Count-MORL\texttt{Count-MORL}. Our method utilizes the count estimates of state-action pairs to quantify model estimation error, marking the first algorithm of demonstrating the efficacy of count-based conservatism in model-based offline deep RL to the best of our knowledge. For our proposed method, we first show that the estimation error is inversely proportional to the frequency of state-action pairs. Secondly, we demonstrate that the learned policy under the count-based conservative model offers near-optimality performance guarantees. Through extensive numerical experiments, we validate that Count-MORL\texttt{Count-MORL} with hash code implementation significantly outperforms existing offline RL algorithms on the D4RL benchmark datasets. The code is accessible at \href\href{https://github.com/oh-lab/Count-MORL}{https://github.com/oh-lab/Count-MORL}.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers8

Ask how each one uses it

Builds on22

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines