Adversarial Semi-Bandits with Moving Arms
Zhiming Huang, Jianping Pan
Abstract
This paper studies a novel multi-agent combinatorial bandit problem called moving semi-bandits involvingagents andarms, extending the problem of semi-bandits with adversar-ial rewards and stochastic arm availabilities (sleeping semi-bandits). The arms move across agents, making each arm available to at most one agent at a time, and the set of available arms for each agent changes over time. In each round, each agent plays up toarms from their own available arm set simultaneously and observes the random loss for each played arm (i.e., semi-bandit feedback). The loss of each arm has no stochastic assumptions, and different agents may generate different random losses for each arm. The primary goal is to minimize the cumulative loss for all agents through collaboration. This bandit problem is motivated by real-world applications, such as traffic scheduling in wireless networks with multiple access points and task assignment for multiple crowdsourcing platforms. To address this challenge, we propose an efficient framework called Moving-FTPL, which guarantees a regret bound ofoverrounds. Moving-Ftplcan reduce the total regret of all K agents by a factor ofcompared to scenarios where agents do not collaborate. Additionally, Moving-FTPL takes a step forward for the long-standing problems of a tighter regret bound for sleeping semi-bandits by significantly improving the state-of-the-art regret bound by a factor ofand imnroving the bound for sleeping adversarial bandits by a factor of. Furth ermore, we showcase a crowdsourcing application to demonstrate the effectiveness of our proposed algorithm when compared with others.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 5f71e910-884c-49a6-8988-7f0067bbe350Related papers
- Bridging the Regret Gap in Combinatorial Thompson Sampling: Worst-Case Guarantees and Algorithmic RefinementZhiming Huang, Bingshan Hu, Jianping PanINFOCOM 2026 · 1 citation
- Follow-the-Perturbed-Leader Nearly Achieves Best-of-Both-Worlds for the m-Set Semi-Bandit ProblemsJingxin Zhan, Yuchen Xin, Chenjie Sun, Zhihua ZhangNeurIPS 2025 · 1 citation
- Adversarial Combinatorial Bandits with Switching Cost and Arm Selection ConstraintsYin Huang, Qingsong Liu, Jie XuINFOCOM 2024 · 10 citations
- A Near-optimal, Scalable and Parallelizable Framework for Stochastic Bandits Robust to Adversarial Corruptions and BeyondZicheng Hu, Cheng ChenNeurIPS 2025
- A Near-Optimal Best-of-Both-Worlds Algorithm for Federated BanditsZicheng Hu, Zihao Wang, Cheng ChenICLR 2026 · 18 citations
