Lune

ICML2022顶会

Contextual Information-Directed Sampling

Botao Hao, Tor Lattimore, Chao Qin

2022年份
19被引次数
8顶会引用

摘要

Information-directed sampling (IDS) has recently demonstrated its potential as a dataefficient reinforcement learning algorithm (Lu et al., 2021) . However, it is still unclear what is the right form of information ratio to optimize when contextual information is available. We investigate the IDS design through two contextual bandit problems: contextual bandits with graph feedback and sparse linear contextual bandits. We provably demonstrate the advantage of contextual IDS over conditional IDS and emphasize the importance of considering the context distribution. The main message is that an intelligent agent should invest more on the actions that are beneficial for the future unseen contexts while the conditional IDS can be myopic. We further propose a computationallyefficient version of contextual IDS based on Actor-Critic and evaluate it empirically on a neural network contextual bandit.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext a90ebcdf-afd8-4bd9-9175-e70fd85453f2

引用它的顶会 Paper8

问问它们各自怎么用它

它引用的顶会 Paper7

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖