BCORLE(λ): An Offline Reinforcement Learning and Evaluation Framework for Coupons Allocation in E-commerce Market
Yang Zhang, Bo Tang, Qingyu Yang, Dou An, Hongyin Tang, Chenyang Xi, Xueying Li, Feiyu Xiong
摘要
Coupons allocation is an important tool for enterprises to increase the activity and loyalty of users on the e-commerce market. One fundamental problem related is how to allocate coupons within a fixed budget while maximizing users' retention on the e-commerce platform. The online e-commerce environment is complicated and ever changing, so it requires the coupons allocation policy learning can quickly adapt to the changes of the company's business strategy. Unfortunately, existing studies with a huge computation overhead can hardly satisfy the requirements of real-time and fast-response in the real world. Specifically, the problem of coupons allocation within a fixed budget is usually formulated as a Lagrangian problem. Existing solutions need to re-learn the policy once the value of Lagrangian multiplier variable λ is updated, causing a great computation overhead. Besides, a mature e-commerce market often faces tens of millions of users and dozens of types of coupons which construct the huge policy space, further increasing the difficulty of solving the problem. To tackle with above problems, we propose a budget constrained offline reinforcement learning and evaluation with λ-generalization (BCORLE(λ)) framework. The proposed method can help enterprises develop a coupons allocation policy which greatly improves users' retention rate on the platform while ensuring the cost does not exceed the budget. Specifically, λ-generalization method is proposed to lead the policy learning process can be executed according to different λ values adaptively, avoiding re-learning new polices from scratch. Thus the computation overhead is greatly reduced. Further, a novel offline reinforcement learning method and an off-policy evaluation algorithm are proposed for policy learning and policy evaluation, respectively. Finally, experiments on the simulation platform and real-world e-commerce market validate the effectiveness of our approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Direct Heterogeneous Causal Learning for Resource Allocation Problems in MarketingHao Zhou, Shaoming Li, Guibin Jiang, Jiaqi Zheng 等AAAI 2023 · 被引用 35 次
- RL-MPCA: A Reinforcement Learning Based Multi-Phase Computation Allocation Approach for Recommender SystemsJiahong Zhou, Shunhui Mao, Guoliang Yang, Bo Tang 等WWW 2023 · 被引用 10 次
- Optimizing Marketing Subsidies via Counterfactual Learning with Asymmetric Reward FunctionXiang Li, Yanghao Xiao, Chunyuan Zheng, Qian Zou 等SIGIR 2026 · 被引用 1 次
- Improve ROI with Causal Learning and Conformal PredictionMeng Ai, Zhuo Chen, Jibin Wang, Jing Shang 等ICDE 2024
它引用的顶会 Paper10
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 被引用 568 次
- Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision ProcessesDongsheng Ding, Kaiqing Zhang, Tamer Basar, Mihailo R. JovanovicNeurIPS 2020 · 被引用 252 次
- Safe Reinforcement Learning in Constrained Markov Decision ProcessesAkifumi Wachi, Yanan SuiICML 2020 · 被引用 190 次
相关 Paper
- Off-Policy Learning with Limited SupplyKoichi Tanaka, Ren Kishimoto, Bushun Kawagishi, Yusuke Narita 等WWW 2026
- SACO: Sequence-Aware Constrained Optimization Framework for Coupon Distribution in E-commerceLi Kong, Bingzhe Wang, Zhou Chen, Suhan Hu 等AAAI 2026
- Coupon Design in Advertising SystemsWeiran Shen, Pingzhong Tang, Xun Wang, Yadong Xu 等AAAI 2021 · 被引用 5 次
- Dynamic Knapsack Optimization Towards Efficient Multi-Channel Sequential AdvertisingXiaotian Hao, Zhaoqing Peng, Yi Ma, Guan Wang 等ICML 2020 · 被引用 29 次
- COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction EstimationJongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz, Nicolas Heess 等ICLR 2022 · 被引用 84 次
