Leveraging Domain Knowledge for Robust Deep Reinforcement Learning in Networking
Ying Zheng, Haoyu Chen, Qingyang Duan, Lixiang Lin, Yiyang Shao, Wei Wang, Xin Wang, Yuedong Xu
摘要
The past few years has witnessed a surge of interest towards deep reinforcement learning (Deep RL) in computer networks. With extraordinary ability of feature extraction, Deep RL has the potential to re-engineer the fundamental resource allocation problems in networking without relying on pre-programmed models or assumptions about dynamic environments. However, such black-box systems suffer from poor robustness, showing high performance variance and poor tail performance. In this work, we propose a unified Teacher-Student learning framework that harnesses rich domain knowledge to improve robustness. The domain-specific algorithms, less performant but more trustable than Deep RL, play the role of teachers providing advice at critical states; the student neural network is steered to maximize the expected reward as usual and mimic the teacher's advice meanwhile. The Teacher-Student method comprises of three modules where the confidence check module locates wrong decisions and risky decisions, the reward shaping module designs a new updating function to incentive the learning of student network, and the prioritized experience replay module to effectively utilize the advised actions. We further implement our Teacher-Student framework in existing video streaming (Pensieve), load balancing (DeepLB) and TCP congestion control (Aurora). Experimental results manifest that the proposed approach reduces the performance standard deviation of DeepLB by 37%; it improves the 90th, 95th and 99th tail performance of Pensieve by 7.6%, 8.8%, 10.7% respectively; and it accelerates the rate of growth of Aurora by 2x at the initial stage, and achieves a more stable performance in dynamic environments.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Eagle: Refining Congestion Control by Learning from the ExpertsSalma Emara, Baochun Li, Yanjiao ChenINFOCOM 2020 · 被引用 65 次
- Genet: automatic curriculum generation for learning adaptation in networkingZhengxu Xia, Yajie Zhou, Francis Y. Yan, Junchen JiangSIGCOMM 2022 · 被引用 57 次
- Stick: A Harmonious Fusion of Buffer-based and Learning-based Approach for Adaptive StreamingTianchi Huang, Chao Zhou, Rui-Xiao Zhang, Chenglei Wu 等INFOCOM 2020 · 被引用 59 次
- Verifying learning-augmented systemsTomer Eliyahu, Yafim Kazak, Guy Katz, Michael SchapiraSIGCOMM 2021 · 被引用 45 次
- Learning Buffer Management Policies for Shared Memory SwitchesMowei Wang, Sijiang Huang, Yong Cui, Wendong Wang 等INFOCOM 2022 · 被引用 12 次
