Online Learning for Load Balancing of Unknown Monotone Resource Allocation Games
Ilai Bistritz, Nicholas Bambos
Abstract
Consider N players that each uses a mixture of K resources. Each of the players' reward functions includes a linear pricing term for each resource that is controlled by the game manager. We assume that the game is strongly monotone, so if each player runs gradient descent, the dynamics converge to a unique Nash equilibrium (NE). Unfortunately, this NE can be inefficient since the total load on a given resource can be very high. In principle, we can control the total loads by tuning the coefficients of the pricing terms. However, finding pricing coefficients that balance the loads requires knowing the players' reward functions and their action sets. Obtaining this game structure information is infeasible in a large-scale network and violates the users' privacy. To overcome this, we propose a simple algorithm that learns to shift the NE of the game to meet the total load constraints by adjusting the pricing coefficients in an online manner. Our algorithm only requires the total load per resource as feedback and does not need to know the reward functions or the action sets. We prove that our algorithm guarantees convergence in L2 to a NE that meets target total load constraints. Simulations show the effectiveness of our approach when applied to smart grid demand-side management or power control in wireless networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Implicit Learning Dynamics in Stackelberg Games: Equilibria Characterization, Convergence Analysis, and Empirical StudyTanner Fiez, Benjamin Chasnov, Lillian J. RatliffICML 2020 · 144 citations
- Dual Mirror Descent for Online Allocation ProblemsSantiago R. Balseiro, Haihao Lu, Vahab S. MirrokniICML 2020 · 102 citations
- My Fair Bandit: Distributed Learning of Max-Min Fairness with Multi-player BanditsIlai Bistritz, Tavor Z. Baharav, Amir Leshem, Nicholas BambosICML 2020 · 40 citations
- Cooperative Multi-player Bandit OptimizationIlai Bistritz, Nicholas BambosNeurIPS 2020 · 31 citations
- Optimally Deceiving a Learning Leader in Stackelberg GamesGeorgios Birmpas, Jiarui Gan, Alexandros Hollender, Francisco J. Marmolejo Cossío et al.NeurIPS 2020 · 25 citations
Related papers
- Nash Equilibria in Games with Playerwise Concave Coupling Constraints: Existence and ComputationPhilip Jordan, Maryam KamgarpourICML 2026
- Online Performative Gradient Descent for Learning Nash Equilibria in Decision-Dependent GamesZihan Zhu, Ethan X. Fang, Zhuoran YangNeurIPS 2023 · 5 citations
- Multi-Agent Distributed Reinforcement Learning for Making Decentralized Offloading DecisionsJing Tan, Ramin Khalili, Holger Karl, Artur HeckerINFOCOM 2022 · 33 citations
- Online Submodular Resource Allocation with Applications to Rebalancing Shared Mobility SystemsPier Giuseppe Sessa, Ilija Bogunovic, Andreas Krause, Maryam KamgarpourICML 2021 · 3 citations
- Single-agent Poisoning Attacks Suffice to Ruin Multi-Agent LearningFan Yao, Yuwei Cheng, Ermin Wei, Haifeng XuICLR 2025
