Progressive Learning with Human Feedback for Personalized Adaptive Video Streaming
Zhaohui Jiang, Xuening Feng, Tianchi Huang, Ruixiao Zhang, Paul Weng, Yifei Zhu
Abstract
Existing quality of experience (QoE)-driven adaptive bitrate (ABR) algorithms either fail to consider personalized QoE or rely on over-simplified QoE models, all resulting in unsatisfactory streaming experiences. Recognizing the wide existence of user feedback schemes in existing streaming applications, we introduce Q+, a framework leveraging progressively gathered personal user opinion scores from multiple interaction sessions for enhanced user-system alignment. Q+ first innovates QoE modeling by incorporating both pairwise ordinal and cardinal preferences constructed from scores. The capturing of both preferences ensures reliable and robust preference representation. Moreover, we design a monotonic neural network as the QoE model to capture the inherent monotonicity property in ABR services, improving model expressivity and generalization ability even with limited human feedback. To align the policy with the progressively updated QoE, we then develop a value-based reinforcement learning (RL) algorithm for bitrate control that integrates reward relabeling and calibrated prioritized experience replay. Extensive experiments reveal that Q+ consistently surpasses state-of-the-art rule-based, control-based, and RL-based baselines within only three sessions, improving QoE by 5.69% to 29.39% across diverse network conditions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 36aa1298-18d2-46ca-b64f-c694e10df8c1Builds on15
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 963 citations
- Defining and Characterizing Reward GamingJoar Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, David KruegerNeurIPS 2022 · 466 citations
- PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-trainingKimin Lee, Laura M. Smith, Pieter AbbeelICML 2021 · 380 citations
- Learning in situ: a randomized experiment in video streamingFrancis Y. Yan, Hudson Ayers, Chenzhi Zhu, Sadjad Fouladi et al.NSDI 2020 · 360 citations
- Bigger, Better, Faster: Human-level Atari with human-level efficiencyMax Schwarzer, Johan S. Obando-Ceron, Aaron C. Courville, Marc G. Bellemare et al.ICML 2023 · 155 citations
Related papers
- Optimizing Adaptive Video Streaming with Human FeedbackTianchi Huang, Rui-Xiao Zhang, Chenglei Wu, Lifeng SunACM MM 2023 · 26 citations
- AraLive: Automatic Reward Adaption for Learning-based Live Video StreamingHuanhuan Zhang, Liu zhuo, Haotian Li, Anfu Zhou et al.ACM MM 2024 · 6 citations
- Improving Generalization for Neural Adaptive Video Streaming via Meta Reinforcement LearningNuowen Kan, Yuankun Jiang, Chenglin Li, Wenrui Dai et al.ACM MM 2022 · 48 citations
- Adaptive Bitrate with User-level QoE Preference for Video StreamingXutong Zuo, Jiayu Yang, Mowei Wang, Yong CuiINFOCOM 2022 · 65 citations
- Buffer Awareness Neural Adaptive Video Streaming for Avoiding Extra Buffer ConsumptionTianchi Huang, Chao Zhou, Rui-Xiao Zhang, Chenglei Wu et al.INFOCOM 2023 · 27 citations
