Two-Stage Constrained Actor-Critic for Short Video Recommendation
Qingpeng Cai, Zhenghai Xue, Chi Zhang, Wanqi Xue, Shuchang Liu, Ruohan Zhan, Xueliang Wang, Tianyou Zuo, Wentao Xie, Dong Zheng, Peng Jiang, Kun Gai
摘要
The wide popularity of short videos on social media poses new opportunities and challenges to optimize recommender systems on the video-sharing platforms. Users sequentially interact with the system and provide complex and multi-faceted responses, including WatchTime and various types of interactions with multiple videos. On the one hand, the platforms aim at optimizing the users’ cumulative WatchTime (main goal) in the long term, which can be effectively optimized by Reinforcement Learning. On the other hand, the platforms also need to satisfy the constraint of accommodating the responses of multiple user interactions (auxiliary goals) such as Like, Follow, Share, etc. In this paper, we formulate the problem of short video recommendation as a Constrained Markov Decision Process (CMDP). We find that traditional constrained reinforcement learning algorithms fail to work well in this setting. We propose a novel two-stage constrained actor-critic method: At stage one, we learn individual policies to optimize each auxiliary signal. In stage two, we learn a policy to (i) optimize the main signal and (ii) stay close to policies learned in the first stage, which effectively guarantees the performance of this main policy on the auxiliaries. Through extensive offline evaluations, we demonstrate the effectiveness of our method over alternatives in both optimizing the main goal as well as balancing the others. We further show the advantage of our method in live experiments of short video recommendations, where it significantly outperforms other baselines in terms of both WatchTime and interactions. Our approach has been fully launched in the production system to optimize user experiences on the platform.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Alleviating Matthew Effect of Offline Reinforcement Learning in Interactive RecommendationChongming Gao, Kexin Huang, Jiawei Chen, Yuan Zhang 等SIGIR 2023 · 被引用 65 次
- PrefRec: Recommender Systems with Human Preferences for Reinforcing Long-term User EngagementWanqi Xue, Qingpeng Cai, Zhenghai Xue, Shuo Sun 等KDD 2023 · 被引用 24 次
- Sequential Recommendation for Optimizing Both Immediate Feedback and Long-term RetentionZiru Liu, Shuchang Liu, Zijian Zhang, Qingpeng Cai 等SIGIR 2024 · 被引用 23 次
- Finite-Time Convergence and Sample Complexity of Multi-Agent Actor-Critic Reinforcement Learning with Average RewardHairi, Jia Liu, Songtao LuICLR 2022 · 被引用 21 次
- Generative Flow Network for Listwise RecommendationShuchang Liu, Qingpeng Cai, Zhankui He, Bowen Sun 等KDD 2023 · 被引用 15 次
它引用的顶会 Paper2
相关 Paper
- MDP2 Forest: A Constrained Continuous Multi-dimensional Policy Optimization Approach for Short-video RecommendationSizhe Yu, Ziyi Liu, Shixiang Wan, Jia Zheng 等KDD 2022 · 被引用 4 次
- Towards End-to-End Alignment of User Satisfaction via Questionnaire in Video RecommendationNa Li, Jiaqi Yu, Minzhi Xie, Tiantian He 等SIGIR 2026
- MTRec: Learning to Align with User Preferences via Mental Reward ModelsMengchen Zhao, Yifan Gao, Yaqing Hou, Xiangyang Li 等NeurIPS 2025 · 被引用 1 次
- Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function ApproximationToshinori Kitamura, Arnob Ghosh, Tadashi Kozuno, Wataru Kumagai 等NeurIPS 2025 · 被引用 5 次
- Exploiting Fine-Grained Skip Behaviors for Micro-Video RecommendationSanghyuck Lee, Sangkeun Park, Jaesung LeeAAAI 2025 · 被引用 2 次
