Drift Plus Optimistic Penalty - A Learning Framework for Stochastic Network Optimization
Sathwik Chadaga, Eytan H. Modiano
Abstract
We consider the problem of joint routing and scheduling in queueing networks, where the edge transmission costs are unknown. At each time-slot, the network controller receives noisy observations of transmission costs only for those edges it selects for transmission. The network controller's objective is to make routing and scheduling decisions so that the total expected cost is minimized. This problem exhibits an explorationexploitation trade-off, however, previous bandit-style solutions cannot be directly applied to this problem due to the queueing dynamics. In order to ensure network stability, the network controller needs to optimize throughput and cost simultaneously. We show that the best achievable cost is lower bounded by the solution to a static optimization problem, and develop a network control policy using techniques from Lyapunov drift-plus-penalty optimization and multi-arm bandits. We show that the policy achieves a sub-linear regret of order O( √ T log T ), as compared to the best policy that has complete knowledge of arrivals and costs. Finally, we evaluate the proposed policy using simulations and show that its regret is indeed sub-linear.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ee4bbcdb-fae6-4ae4-9f8e-8fab75681363Builds on1
Related papers
- A Lyapunov-Based Methodology for Constrained Optimization with Bandit FeedbackSemih Cayci, Yilin Zheng, Atilla EryilmazAAAI 2022 · 12 citations
- Learning-based Scheduling for Information Gathering with QoS ConstraintsQingsong Liu, Weihang Xu, Zhixuan FangINFOCOM 2024 · 5 citations
- Geometric Exploration for Online ControlOrestis Plevrakis, Elad HazanNeurIPS 2020 · 12 citations
- Faster Convergence for Unknown-Game BanditsZhiming Huang, Jianping PanINFOCOM 2025
- Online Packet Scheduling with Deadlines and LearningGianmarco Genalti, Achraf Azize, Vianney PerchetICML 2026
