MDP2 Forest: A Constrained Continuous Multi-dimensional Policy Optimization Approach for Short-video Recommendation
Sizhe Yu, Ziyi Liu, Shixiang Wan, Jia Zheng, Zang Li, Fan Zhou
摘要
In the ecology of short video platforms, the optimal exposure proportion of each video category is crucial to guide recommendation systems and content production in a macroscopic way. Though extensive studies on recommendation systems are devoted to providing the most well-matched videos for each view request, fitting the data without considering inherent biases such as selection bias and exposure bias will result in serious issues. In this paper, we formalize the exposure proportion strategy as a policy-making problem with multi-dimensional continuous treatment under certain constraints from a causal inference point of view. We propose a novel ensemble policy learning method based on causal trees, called Maximum Difference of Preference Point Forest (MDP2 Forest), which overcomes the shortcomings of existing policy learning approaches. Experimental results on both simulated and synthetic datasets show the superiority of our algorithm compared to other policy learning or causal inference methods in terms of the treatment estimation accuracy and the mean regret. Furthermore, the proposed MDP2 Forest method can also adapt to a wide range of business settings such as imposing different kinds of constraints on the multi-dimensional treatment.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Uplift Modeling for Target User Attacks on Recommender SystemsWenjie Wang, Changsheng Wang, Fuli Feng, Wentao Shi 等WWW 2024 · 被引用 11 次
- Treatment Effect Estimation for User Interest Exploration on Recommender SystemsJiaju Chen, Wenjie Wang, Chongming Gao, Peng Wu 等SIGIR 2024 · 被引用 8 次
相关 Paper
- Two-Stage Constrained Actor-Critic for Short Video RecommendationQingpeng Cai, Zhenghai Xue, Chi Zhang, Wanqi Xue 等WWW 2023 · 被引用 60 次
- Full Stage Learning to Rank: A Unified Framework for Multi-Stage SystemsKai Zheng, Haijun Zhao, Rui Huang, Beichuan Zhang 等WWW 2024 · 被引用 24 次
- Scalable Multi-Action Offline Policy Learning with an m-ary TreeShusei EshimaKDD 2026
- LBCF: A Large-Scale Budget-Constrained Causal Forest AlgorithmMeng Ai, Biao Li, Heyang Gong, Qingwei Yu 等WWW 2022 · 被引用 27 次
- CORAL: Uncertainty-Aware Regulation of Exposure Concentration in Recommender SystemsNitin Bisht, Linjiang Guo, Xiuwen Gong, Huan Huo 等ICML 2026
