Progressive Learning with Human Feedback for Personalized Adaptive Video Streaming
Zhaohui Jiang, Xuening Feng, Tianchi Huang, Ruixiao Zhang, Paul Weng, Yifei Zhu
摘要
Existing quality of experience (QoE)-driven adaptive bitrate (ABR) algorithms either fail to consider personalized QoE or rely on over-simplified QoE models, all resulting in unsatisfactory streaming experiences. Recognizing the wide existence of user feedback schemes in existing streaming applications, we introduce Q+, a framework leveraging progressively gathered personal user opinion scores from multiple interaction sessions for enhanced user-system alignment. Q+ first innovates QoE modeling by incorporating both pairwise ordinal and cardinal preferences constructed from scores. The capturing of both preferences ensures reliable and robust preference representation. Moreover, we design a monotonic neural network as the QoE model to capture the inherent monotonicity property in ABR services, improving model expressivity and generalization ability even with limited human feedback. To align the policy with the progressively updated QoE, we then develop a value-based reinforcement learning (RL) algorithm for bitrate control that integrates reward relabeling and calibrated prioritized experience replay. Extensive experiments reveal that Q+ consistently surpasses state-of-the-art rule-based, control-based, and RL-based baselines within only three sessions, improving QoE by 5.69% to 29.39% across diverse network conditions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- Defining and Characterizing Reward GamingJoar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, David KruegerNeurIPS 2022 · 被引用 466 次
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 被引用 380 次
- Learning in situ: a randomized experiment in video streamingFrancis Y. Yan, Hudson Ayers, Chenzhi Zhu, Sadjad Fouladi 等NSDI 2020 · 被引用 360 次
- Bigger, Better, Faster: Human-level Atari with human-level efficiencyMax Schwarzer, Johan S. Obando-Ceron, Aaron C. Courville, Marc G. Bellemare 等ICML 2023 · 被引用 155 次
相关 Paper
- Optimizing Adaptive Video Streaming with Human FeedbackTianchi Huang, Rui-Xiao Zhang, Chenglei Wu, Lifeng SunACM MM 2023 · 被引用 26 次
- AraLive: Automatic Reward Adaption for Learning-based Live Video StreamingHuanhuan Zhang, Liu zhuo, Haotian Li, Anfu Zhou 等ACM MM 2024 · 被引用 6 次
- Improving Generalization for Neural Adaptive Video Streaming via Meta Reinforcement LearningNuowen Kan, Yuankun Jiang, Chenglin Li, Wenrui Dai 等ACM MM 2022 · 被引用 48 次
- Adaptive Bitrate with User-level QoE Preference for Video StreamingXutong Zuo, Jiayu Yang, Mowei Wang, Yong CuiINFOCOM 2022 · 被引用 65 次
- Buffer Awareness Neural Adaptive Video Streaming for Avoiding Extra Buffer ConsumptionTianchi Huang, Chao Zhou, Rui-Xiao Zhang, Chenglei Wu 等INFOCOM 2023 · 被引用 27 次
