Two-Stage Constrained Actor-Critic for Short Video Recommendation
Qingpeng Cai, Zhenghai Xue, Chi Zhang, Wanqi Xue, Shuchang Liu, Ruohan Zhan, Xueliang Wang, Tianyou Zuo, Wentao Xie, Dong Zheng, Peng Jiang, Kun Gai
Abstract
The wide popularity of short videos on social media poses new opportunities and challenges to optimize recommender systems on the video-sharing platforms. Users sequentially interact with the system and provide complex and multi-faceted responses, including WatchTime and various types of interactions with multiple videos. On the one hand, the platforms aim at optimizing the users’ cumulative WatchTime (main goal) in the long term, which can be effectively optimized by Reinforcement Learning. On the other hand, the platforms also need to satisfy the constraint of accommodating the responses of multiple user interactions (auxiliary goals) such as Like, Follow, Share, etc. In this paper, we formulate the problem of short video recommendation as a Constrained Markov Decision Process (CMDP). We find that traditional constrained reinforcement learning algorithms fail to work well in this setting. We propose a novel two-stage constrained actor-critic method: At stage one, we learn individual policies to optimize each auxiliary signal. In stage two, we learn a policy to (i) optimize the main signal and (ii) stay close to policies learned in the first stage, which effectively guarantees the performance of this main policy on the auxiliaries. Through extensive offline evaluations, we demonstrate the effectiveness of our method over alternatives in both optimizing the main goal as well as balancing the others. We further show the advantage of our method in live experiments of short video recommendations, where it significantly outperforms other baselines in terms of both WatchTime and interactions. Our approach has been fully launched in the production system to optimize user experiences on the platform.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 058b3534-a91d-4f0e-bdb4-f18dbbe48d7dCited by top-tier papers12
- Alleviating Matthew Effect of Offline Reinforcement Learning in Interactive RecommendationChongming Gao, Kexin Huang, Jiawei Chen, Yuan Zhang et al.SIGIR 2023 · 65 citations
- PrefRec: Recommender Systems with Human Preferences for Reinforcing Long-term User EngagementWanqi Xue, Qingpeng Cai, Zhenghai Xue, Shuo Sun et al.KDD 2023 · 24 citations
- Sequential Recommendation for Optimizing Both Immediate Feedback and Long-term RetentionZiru Liu, Shuchang Liu, Zijian Zhang, Qingpeng Cai et al.SIGIR 2024 · 23 citations
- Finite-Time Convergence and Sample Complexity of Multi-Agent Actor-Critic Reinforcement Learning with Average RewardHairi, Jia Liu, Songtao LuICLR 2022 · 21 citations
- Generative Flow Network for Listwise RecommendationShuchang Liu, Qingpeng Cai, Zhankui He, Bowen Sun et al.KDD 2023 · 15 citations
Builds on2
Related papers
- MDP2 Forest: A Constrained Continuous Multi-dimensional Policy Optimization Approach for Short-video RecommendationSizhe Yu, Ziyi Liu, Shixiang Wan, Jia Zheng et al.KDD 2022 · 4 citations
- Towards End-to-End Alignment of User Satisfaction via Questionnaire in Video RecommendationNa Li, Jiaqi Yu, Minzhi Xie, Tiantian He et al.SIGIR 2026
- MTRec: Learning to Align with User Preferences via Mental Reward ModelsMengchen Zhao, Yifan Gao, Yaqing Hou, Xiangyang Li et al.NeurIPS 2025 · 1 citation
- Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function ApproximationToshinori Kitamura, Arnob Ghosh, Tadashi Kozuno, Wataru Kumagai et al.NeurIPS 2025 · 5 citations
- Exploiting Fine-Grained Skip Behaviors for Micro-Video RecommendationSanghyuck Lee, Sangkeun Park, Jaesung LeeAAAI 2025 · 2 citations
