MDP2 Forest: A Constrained Continuous Multi-dimensional Policy Optimization Approach for Short-video Recommendation
Sizhe Yu, Ziyi Liu, Shixiang Wan, Jia Zheng, Zang Li, Fan Zhou
Abstract
In the ecology of short video platforms, the optimal exposure proportion of each video category is crucial to guide recommendation systems and content production in a macroscopic way. Though extensive studies on recommendation systems are devoted to providing the most well-matched videos for each view request, fitting the data without considering inherent biases such as selection bias and exposure bias will result in serious issues. In this paper, we formalize the exposure proportion strategy as a policy-making problem with multi-dimensional continuous treatment under certain constraints from a causal inference point of view. We propose a novel ensemble policy learning method based on causal trees, called Maximum Difference of Preference Point Forest (MDP2 Forest), which overcomes the shortcomings of existing policy learning approaches. Experimental results on both simulated and synthetic datasets show the superiority of our algorithm compared to other policy learning or causal inference methods in terms of the treatment estimation accuracy and the mean regret. Furthermore, the proposed MDP2 Forest method can also adapt to a wide range of business settings such as imposing different kinds of constraints on the multi-dimensional treatment.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get bc820622-6d42-49d7-a553-60b7fde008f9Cited by top-tier papers2
- Uplift Modeling for Target User Attacks on Recommender SystemsWenjie Wang, Changsheng Wang, Fuli Feng, Wentao Shi et al.WWW 2024 · 11 citations
- Treatment Effect Estimation for User Interest Exploration on Recommender SystemsJiaju Chen, Wenjie Wang, Chongming Gao, Peng Wu et al.SIGIR 2024 · 8 citations
Related papers
- Two-Stage Constrained Actor-Critic for Short Video RecommendationQingpeng Cai, Zhenghai Xue, Chi Zhang, Wanqi Xue et al.WWW 2023 · 60 citations
- Full Stage Learning to Rank: A Unified Framework for Multi-Stage SystemsKai Zheng, Haijun Zhao, Rui Huang, Beichuan Zhang et al.WWW 2024 · 24 citations
- Scalable Multi-Action Offline Policy Learning with an m-ary TreeShusei EshimaKDD 2026
- LBCF: A Large-Scale Budget-Constrained Causal Forest AlgorithmMeng Ai, Biao Li, Heyang Gong, Qingwei Yu et al.WWW 2022 · 27 citations
- CORAL: Uncertainty-Aware Regulation of Exposure Concentration in Recommender SystemsNitin Bisht, Linjiang Guo, Xiuwen Gong, Huan Huo et al.ICML 2026
