Winning the L2RPN Challenge: Power Grid Management via Semi-Markov Afterstate Actor-Critic
Deunsol Yoon, Sunghoon Hong, Byung-Jun Lee, Kee-Eung Kim
Abstract
Safe and reliable electricity transmission in power grids is crucial for modern society. It is thus quite natural that there has been a growing interest in the automatic management of power grids, exemplified by the Learning to Run a Power Network Challenge (L2RPN), modeling the problem as a reinforcement learning (RL) task. However, it is highly challenging to manage a real-world scale power grid, mostly due to the massive scale of its state and action space. In this paper, we present an off-policy actor-critic approach that effectively tackles the unique challenges in power grid management by RL, adopting the hierarchical policy together with the afterstate representation. Our agent ranked first in the latest challenge (L2RPN WCCI 2020), being able to avoid disastrous situations while maintaining the highest level of operational efficiency in every test scenarios. This paper provides a formal description of the algorithmic aspect of our approach, as well as further experimental studies on diverse power grids.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers2
- Stabilizing Voltage in Power Distribution Networks via Multi-Agent Reinforcement Learning with TransformerMinrui Wang, Mingxiao Feng, Wengang Zhou, Houqiang LiKDD 2022 · 14 citations
- Think Twice, Act Once: A Co-Evolution Framework of LLM and RL for Large-Scale Decision MakingXu Wan, Wenyue Xu, Chao Yang, Mingyang SunICML 2025
Related papers
- MARL2Grid-TR: A Multi-Agent RL Benchmark in Power Grid OperationsEnrico Marchesini, Eva Boguslawski, Alessandro Leite, Christopher Amato et al.ICLR 2026
- Graph Reinforcement Learning for Network Control via Bi-Level OptimizationDaniele Gammelli, James Harrison, Kaidi Yang, Marco Pavone et al.ICML 2023 · 14 citations
- A Hierarchical Reinforcement Learning Based Optimization Framework for Large-scale Dynamic Pickup and Delivery ProblemsYi Ma, Xiaotian Hao, Jianye Hao, Jiawen Lu et al.NeurIPS 2021 · 100 citations
- Causality Meets Locality: Provably Generalizable and Scalable Policy Learning for Networked SystemsHao Liang, Shuqing Shi, Yudi Zhang, Biwei Huang et al.NeurIPS 2025 · 1 citation
- Multi-Agent Reinforcement Learning for Active Voltage Control on Power Distribution NetworksJianhong Wang, Wangkun Xu, Yunjie Gu, Wenbin Song et al.NeurIPS 2021 · 216 citations
