OnRL: improving mobile video telephony via online reinforcement learning
Huanhuan Zhang, Anfu Zhou, Jiamin Lu, Ruoxuan Ma, Yuhan Hu, Cong Li, Xinyu Zhang, Huadong Ma, Xiaojiang Chen
摘要
Machine learning models, particularly reinforcement learning (RL), have demonstrated great potential in optimizing video streaming applications. However, the state-of-the-art solutions are limited to an "offline learning" paradigm, i.e., the RL models are trained in simulators and then are operated in real networks. As a result, they inevitably suffer from the simulation-to-reality gap, showing far less satisfactory performance under real conditions compared with simulated environment. In this work, we close the gap by proposing OnRL, an online RL framework for real-time mobile video telephony. OnRL puts many individual RL agents directly into the video telephony system, which make video bitrate decisions in real-time and evolve their models over time. OnRL then aggregates these agents to form a high-level RL model that can help each individual to react to unseen network conditions. Moreover, OnRL incorporates novel mechanisms to handle the adverse impacts of inherent video traffic dynamics, and to eliminate risks of quality degradation caused by the RL model's exploration attempts. We implement OnRL on a mainstream operational video telephony system, Alibaba Taobao-live. In a month-long evaluation with 543 hours of video sessions from 151 real-world mobile users, OnRL outperforms the prior algorithms significantly, reducing video stalling rate by 14.22% while maintaining similar video quality.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper18
- LiveNet: a low-latency video transport network for large-scale live streamingJinyang Li, Zhenyu Li, Ri Lu, Kai Xiao 等SIGCOMM 2022 · 被引用 59 次
- GRACE: Loss-Resilient Real-Time Video through Neural CodecsYihua Cheng, Ziyi Zhang, Hanchen Li, Anton Arapin 等NSDI 2024 · 被引用 53 次
- A Workload-Aware DVFS Robust to Concurrent Tasks for Mobile DevicesChengdong Lin, Kun Wang, Zhenjiang Li, Yu PuMobiCom 2023 · 被引用 52 次
- Enabling High Quality Real-Time Communications with Adaptive Frame-RateZili Meng, Tingfeng Wang, Yixin Shen, Bo Wang 等NSDI 2023 · 被引用 46 次
- Pudica: Toward Near-Zero Queuing Delay in Congestion Control for Cloud GamingShibo Wang, Shusen Yang, Xiao Kong, Chenglei Wu 等NSDI 2024 · 被引用 32 次
相关 Paper
- From Ember to Blaze: Swift Interactive Video Adaptation via Meta-Reinforcement LearningXuedou Xiao, Mingxuan Yan, Yingying Zuo, Boxi Liu 等INFOCOM 2023 · 被引用 14 次
- Exstream: A Delay-minimized Streaming System with Explicit Frame Queueing Delay MeasurementShinik Park, Sanghyun Han, Junseon Kim, Jongyun Lee 等INFOCOM 2024 · 被引用 1 次
- AraLive: Automatic Reward Adaption for Learning-based Live Video StreamingHuanhuan Zhang, Liu zhuo, Haotian Li, Anfu Zhou 等ACM MM 2024 · 被引用 6 次
- EAVS: Edge-assisted Adaptive Video Streaming with Fine-grained Serverless PipelinesBiao Hou, Song Yang, Fernando A. Kuipers, Lei Jiao 等INFOCOM 2023 · 被引用 26 次
- Loki: improving long tail performance of learning-based real-time video adaptation by fusing rule-based modelsHuanhuan Zhang, Anfu Zhou, Yuhan Hu, Chaoyue Li 等MobiCom 2021 · 被引用 78 次
