Lune

ICML2022Top-tier venue

Contextual Information-Directed Sampling

Botao Hao, Tor Lattimore, Chao Qin

2022Year
19Citations
8Top-tier citations

Abstract

Information-directed sampling (IDS) has recently demonstrated its potential as a dataefficient reinforcement learning algorithm (Lu et al., 2021) . However, it is still unclear what is the right form of information ratio to optimize when contextual information is available. We investigate the IDS design through two contextual bandit problems: contextual bandits with graph feedback and sparse linear contextual bandits. We provably demonstrate the advantage of contextual IDS over conditional IDS and emphasize the importance of considering the context distribution. The main message is that an intelligent agent should invest more on the actions that are beneficial for the future unseen contexts while the conditional IDS can be myopic. We further propose a computationallyefficient version of contextual IDS based on Actor-Critic and evaluate it empirically on a neural network contextual bandit.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext a90ebcdf-afd8-4bd9-9175-e70fd85453f2

Cited by top-tier papers8

Ask how each one uses it

Builds on7

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines