Lune

ICML2020顶会

Improved Optimistic Algorithms for Logistic Bandits

Louis Faury, Marc Abeille, Clément Calauzènes, Olivier Fercoq

2020年份
127被引次数
72顶会引用

摘要

The generalized linear bandit framework has attracted a lot of attention in recent years by extending the well-understood linear setting and allowing to model richer reward structures. It notably covers the logistic model, widely used when rewards are binary. For logistic bandits, the frequentist regret guarantees of existing algorithms are O~(κT)\tilde{\mathcal{O}}(\kappa \sqrt{T}), where κ\kappa is a problem-dependent constant. Unfortunately, κ\kappa can be arbitrarily large as it scales exponentially with the size of the decision set. This may lead to significantly loose regret bounds and poor empirical performance. In this work, we study the logistic bandit with a focus on the prohibitive dependencies introduced by κ\kappa. We propose a new optimistic algorithm based on a finer examination of the non-linearities of the reward function. We show that it enjoys a O~(T)\tilde{\mathcal{O}}(\sqrt{T}) regret with no dependency in κ\kappa, but for a second order term. Our analysis is based on a new tail-inequality for self-normalized martingales, of independent interest.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper72

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖