Multi-Task Reinforcement Learning for Collaborative Network Optimization in Data Centers
Ting Wang, Kai Cheng, Xiao Du
Abstract
As data center networks increasingly grow in complexity and scale, efficiently managing traffic scheduling and congestion control becomes crucial for optimizing network performance. Traditional single-task optimization strategies often fall short, failing to adequately address the interplay between different tasks and resulting in suboptimal performance with inefficiencies and robustness issues. To tackle these challenges, this paper proposes a novel Multi-Task Reinforcement Learning (MTRL)-based collaborative Network Optimization scheme, termed MTRLNO, which establishes a structured framework with central and edge systems (i.e., hosts and switches). The SDN-enabled central system incorporates an MTRL agent that simultaneously optimizes traffic scheduling and congestion control tasks, leveraging global network state information to formulate instructive optimization policies for edge systems. Switches implement decentralized multi-agent RL agents to facilitate automatic ECN tuning for congestion control, with the ability to handle incast issues. Hosts feature an MTRL-guided Multiple Level Feedback Queue (MLFQ) demotion threshold adjustment scheme for adaptive traffic scheduling. We further develop a Prioritized Experience Replay-based Soft Actor-Critic (PERSAC) algorithm to enhance learning efficiency and a customized multi-task learning algorithm via improved parameter-sharing to effectively adapt across multiple tasks. Experimental results demonstrate that MTRLNO significantly outperforms state-of-the-art approaches in terms of FCT, latency, and robustness across diverse network conditions.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ad66ee8e-4102-4afd-b632-9856af094f68Related papers
- ACC: automatic ECN tuning for high-speed datacenter networksSiyu Yan, Xiaoliang Wang, Xiaolong Zheng, Yinben Xia et al.SIGCOMM 2021 · 95 citations
- TapFinger: Task Placement and Fine-Grained Resource Allocation for Edge Machine LearningYihong Li, Tianyu Zeng, Xiaoxi Zhang, Jingpu Duan et al.INFOCOM 2023 · 39 citations
- DRL-OR: Deep Reinforcement Learning-based Online Routing for Multi-type Service RequirementsChenyi Liu, Mingwei Xu, Yuan Yang, Nan GengINFOCOM 2021 · 85 citations
- Multi-objective congestion controlYiqing Ma, Han Tian, Xudong Liao, Junxue Zhang et al.EuroSys 2022 · 52 citations
- AUTO: Adaptive Congestion Control Based on Multi-Objective Reinforcement Learning for the Satellite-Ground Integrated NetworkXu Li, Feilong Tang, Jiacheng Liu, Laurence T. Yang et al.USENIX ATC 2021 · 38 citations
