Lune

ICML2023顶会

Model-based Offline Reinforcement Learning with Count-based Conservatism

Byeongchan Kim, Min-hwan Oh

2023年份
19被引次数
8顶会引用

摘要

In this paper, we propose a model-based offline reinforcement learning method that integrates count-based conservatism, named Count-MORL\texttt{Count-MORL}. Our method utilizes the count estimates of state-action pairs to quantify model estimation error, marking the first algorithm of demonstrating the efficacy of count-based conservatism in model-based offline deep RL to the best of our knowledge. For our proposed method, we first show that the estimation error is inversely proportional to the frequency of state-action pairs. Secondly, we demonstrate that the learned policy under the count-based conservative model offers near-optimality performance guarantees. Through extensive numerical experiments, we validate that Count-MORL\texttt{Count-MORL} with hash code implementation significantly outperforms existing offline RL algorithms on the D4RL benchmark datasets. The code is accessible at \href\href{https://github.com/oh-lab/Count-MORL}{https://github.com/oh-lab/Count-MORL}.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper8

问问它们各自怎么用它

它引用的顶会 Paper22

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖