Lune

ICML2021Top-tier venue

Adversarial Dueling Bandits

Aadirupa Saha, Tomer Koren, Yishay Mansour

2021Year
35Citations
21Top-tier citations

Abstract

We introduce the problem of regret minimization in Adversarial Dueling Bandits. As in classic Dueling Bandits, the learner has to repeatedly choose a pair of items and observe only a relative binary `win-loss' feedback for this pair, but here this feedback is generated from an arbitrary preference matrix, possibly chosen adversarially. Our main result is an algorithm whose TT-round regret compared to the Borda-winner from a set of KK items is O~(K1/3T2/3)\tilde{O}(K^{1/3}T^{2/3}), as well as a matching Ω(K1/3T2/3)\Omega(K^{1/3}T^{2/3}) lower bound. We also prove a similar high probability regret bound. We further consider a simpler fixed-gap adversarial setup, which bridges between two extreme preference feedback models for dueling bandits: stationary preferences and an arbitrary sequence of preferences. For the fixed-gap adversarial setup we give an O~((K/Δ2)log⁡T)\smash{ \tilde{O}((K/\Delta^2)\log{T}) } regret algorithm, where Δ\Delta is the gap in Borda scores between the best item and all other items, and show a lower bound of Ω(K/Δ2)\Omega(K/\Delta^2) indicating that our dependence on the main problem parameters KK and Δ\Delta is tight (up to logarithmic factors).

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 678557e8-9931-473b-bbdf-6d1a3f6cc7ea

Cited by top-tier papers21

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines