Graph Reinforcement Learning for Network Control via Bi-Level Optimization
Daniele Gammelli, James Harrison, Kaidi Yang, Marco Pavone, Filipe Rodrigues, Francisco C. Pereira
Abstract
Optimization problems over dynamic networks have been extensively studied and widely used in the past decades to formulate numerous real-world problems. However, (1) traditional optimization-based approaches do not scale to large networks, and (2) the design of good heuristics or approximation algorithms often requires significant manual trial-and-error. In this work, we argue that data-driven strategies can automate this process and learn efficient algorithms without compromising optimality. To do so, we present network control problems through the lens of reinforcement learning and propose a graph network-based framework to handle a broad class of problems. Instead of naively computing actions over high-dimensional graph elements, e.g., edges, we propose a bi-level formulation where we (1) specify a desired next state via RL, and (2) solve a convex program to best achieve it, leading to drastically improved scalability and performance. We further highlight a collection of desirable features to system designers, investigate design decisions, and present experiments on real-world control problems showing the utility, scalability, and flexibility of our framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Hierarchical Decision Making with Structured Policies: A Principled Design via Inverse OptimizationYuexuan Wang, Jingyuan Zhou, Kaidi YangICML 2026
- Offline Hierarchical Reinforcement Learning via Inverse OptimizationCarolin Schmidt, Daniele Gammelli, James Harrison, Marco Pavone et al.ICLR 2025
Builds on4
- What Can Neural Networks Reason About?Keyulu Xu, Jingling Li, Mozhi Zhang, Simon S. Du et al.ICLR 2020 · 281 citations
- Neur2SP: Neural Two-Stage Stochastic ProgrammingRahul Patel, Justin Dumouchelle, Elias B. Khalil, Merve BodurNeurIPS 2022 · 63 citations
- The Differentiable Cross-Entropy MethodBrandon Amos, Denis YaratsICML 2020 · 60 citations
- Why Should I Trust You, Bellman? The Bellman Error is a Poor Replacement for Value ErrorScott Fujimoto, David Meger, Doina Precup, Ofir Nachum et al.ICML 2022 · 43 citations
Related papers
- Reward Propagation Using Graph Convolutional NetworksMartin Klissarov, Doina PrecupNeurIPS 2020 · 28 citations
- Learning to Accelerate Traffic Allocation Over Large-Scale NetworksZhaoxing Yang, Guiyun Fan, Anjie Cao, Yuchen Guo et al.INFOCOM 2025 · 3 citations
- Breaking the Scalability Barrier in Constrained Graph-Based Networked Control via Decision-Focused LearningZhaoxing Yang, Yuchen Guo, Wenlong Li, Guiyun Fan et al.WWW 2026
- Learning to Schedule Learning rate with Graph Neural NetworksYuanhao Xiong, Li-Cheng Lan, Xiangning Chen, Ruochen Wang et al.ICLR 2022 · 18 citations
- Learning Mean Field Control on Sparse GraphsChristian Fabian, Kai Cui, Heinz KoepplICML 2025
