Lune

NeurIPS2022顶会

DASCO: Dual-Generator Adversarial Support Constrained Offline Reinforcement Learning

Quan Vuong, Aviral Kumar, Sergey Levine, Yevgen Chebotar

2022年份
9被引次数
6顶会引用

摘要

In offline RL, constraining the learned policy to remain close to the data is essential 1 to prevent the policy from outputting out-of-distribution (OOD) actions with erro-2 neously overestimated values. In principle, generative adversarial networks (GAN) 3 can provide an elegant solution to do so, with the discriminator directly providing 4 a probability that quantifies distributional shift. However, in practice, GAN-based 5 offline RL methods have not outperformed alternative approaches, perhaps because 6 the generator is trained to both fool the discriminator and maximize return – two 7 objectives that are often at odds with each other. In this paper, we show that the 8 issue of conflicting objectives can be resolved by training two generators: one that 9 maximizes return, with the other capturing the “remainder” of the data distribution 10 in the offline dataset, such that the mixture of the two is close to the behavior policy. 11 We show that not only does having two generators enable an effective GAN-based 12 offline RL method, but also approximates a support constraint, where the policy 13 does not need to match the entire data distribution, but only the slice of the data 14 that leads to high long term performance. We name our method DASCO, for 15 D ual-Generator A dversarial S upport C onstrained O ffline RL. On benchmark tasks 16 that require learning from sub-optimal data, DASCO significantly outperforms 17 prior methods that enforce distribution constraint. 18

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper6

问问它们各自怎么用它

它引用的顶会 Paper13

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖