Identification of the Generalized Condorcet Winner in Multi-dueling Bandits
Björn Haddenhorst, Viktor Bengs, Eyke Hüllermeier
摘要
The reliable identification of the “best” arm while keeping the sample complexity as low as possible is a common task in the field of multi-armed bandits. In the multi-dueling variant of multi-armed bandits, where feedback is provided in the form of a winning arm among as set of k chosen ones, a reasonable notion of best arm is the generalized Condorcet winner (GCW). The latter is an the arm that has the greatest probability of being the winner in each subset containing it. In this paper, we derive lower bounds on the sample complexity for the task of identifying the GCW under various assumptions. As a by-product, our lower bound results provide new insights for the special case of dueling bandits (k = 2). We propose the Dvoretzky–Kiefer–Wolfowitz tournament (DKWT) algorithm, which we prove to be nearly optimal. In a numerical study, we show that DKWT empirically outperforms current state-of-the-art algorithms, even in the special case of dueling bandits or under a Plackett-Luce assumption on the feedback mechanism.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Finding Optimal Arms in Non-stochastic Combinatorial Bandits with Semi-bandit Feedback and Finite BudgetJasmin Brandt, Viktor Bengs, Björn Haddenhorst, Eyke HüllermeierNeurIPS 2022 · 被引用 9 次
- Clustering Items through Bandit Feedback: Finding the Right Feature out of ManyMaximilian Graf, Victor Thuot, Nicolas VerzelenICML 2025
- Achieving Nearly-Optimal Regret and Sample Complexity in Dueling Bandits with Applications in Online RecommendationsLanjihong Ma, Yao-Xiang Ding, Zhen-Yu Zhang, Zhi-Hua ZhouKDD 2025
- Optimal Top- Identification from Pairwise ComparisonsMotti Goldberger, Nils RudiICML 2026
它引用的顶会 Paper5
- Choice BanditsArpit Agarwal, Nicholas Johnson, Shivani AgarwalNeurIPS 2020 · 被引用 19 次
- From PAC to Instance-Optimal Sample Complexity in the Plackett-Luce ModelAadirupa Saha, Aditya GopalanICML 2020 · 被引用 16 次
- The Sample Complexity of Best-k Items Selection from Pairwise ComparisonsWenbo Ren, Jia Liu, Ness B. ShroffICML 2020 · 被引用 14 次
- Preselection BanditsViktor Bengs, Eyke HüllermeierICML 2020 · 被引用 7 次
- Sequential Mode Estimation with Oracle QueriesDhruti Shah, Tuhinangshu Choudhury, Nikhil Karamchandani, Aditya GopalanAAAI 2020 · 被引用 7 次
相关 Paper
- Dueling Bandits with Team ComparisonsLee Cohen, Ulrike Schmidt-Kraepelin, Yishay MansourNeurIPS 2021 · 被引用 1 次
- Combinatorial Pure Exploration for Dueling BanditWei Chen, Yihan Du, Longbo Huang, Haoyu ZhaoICML 2020 · 被引用 14 次
- Batched Dueling BanditsArpit Agarwal, Rohan Ghuge, Viswanath NagarajanICML 2022 · 被引用 12 次
- An Asymptotically Optimal Batched Algorithm for the Dueling Bandit ProblemArpit Agarwal, Rohan Ghuge, Viswanath NagarajanNeurIPS 2022 · 被引用 2 次
- On Weak Regret Analysis for Dueling BanditsEl Mehdi Saad, Alexandra Carpentier, Tomás Kocák, Nicolas VerzelenNeurIPS 2024 · 被引用 5 次
