Lune

ICML2023顶会

Reinforcement Learning from Passive Data via Latent Intentions

Dibya Ghosh, Chethan Anand Bhateja, Sergey Levine

2023年份
69被引次数
40顶会引用

摘要

Passive observational data, such as human videos, is abundant and rich in information, yet remains largely untapped by current RL methods. Perhaps surprisingly, we show that passive data, despite not having reward or action labels, can still be used to learn features that accelerate downstream RL. Our approach learns from passive data by modeling intentions: measuring how the likelihood of future outcomes change when the agent acts to achieve a particular task. We propose a temporal difference learning objective to learn about intentions, resulting in an algorithm similar to conventional RL, but which learns entirely from passive data. When optimizing this objective, our agent simultaneously learns representations of states, of policies, and of possible outcomes in an environment, all from raw observational data. Both theoretically and empirically, this scheme learns features amenable for value prediction for downstream tasks, and our experiments demonstrate the ability to learn from many forms of passive data, including cross-embodiment video data and YouTube videos. Figure 1 : We seek to extract a general knowledge of how an agent may act to influence its environment by pre-training on passive data. Our approach models the effects of acting with intention: jointly learning a latent space of agent intentions and an intention-conditioned value function that estimates the likelihood of witnessing any given outcome in the future when acting according to some latent intention.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 134a1f06-9279-4fcd-83ec-9945a2202f9e

引用它的顶会 Paper40

问问它们各自怎么用它

它引用的顶会 Paper12

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖