BCORLE(λ): An Offline Reinforcement Learning and Evaluation Framework for Coupons Allocation in E-commerce Market
Yang Zhang, Bo Tang, Qingyu Yang, Dou An, Hongyin Tang, Chenyang Xi, Xueying Li, Feiyu Xiong
Abstract
Coupons allocation is an important tool for enterprises to increase the activity and loyalty of users on the e-commerce market. One fundamental problem related is how to allocate coupons within a fixed budget while maximizing users' retention on the e-commerce platform. The online e-commerce environment is complicated and ever changing, so it requires the coupons allocation policy learning can quickly adapt to the changes of the company's business strategy. Unfortunately, existing studies with a huge computation overhead can hardly satisfy the requirements of real-time and fast-response in the real world. Specifically, the problem of coupons allocation within a fixed budget is usually formulated as a Lagrangian problem. Existing solutions need to re-learn the policy once the value of Lagrangian multiplier variable λ is updated, causing a great computation overhead. Besides, a mature e-commerce market often faces tens of millions of users and dozens of types of coupons which construct the huge policy space, further increasing the difficulty of solving the problem. To tackle with above problems, we propose a budget constrained offline reinforcement learning and evaluation with λ-generalization (BCORLE(λ)) framework. The proposed method can help enterprises develop a coupons allocation policy which greatly improves users' retention rate on the platform while ensuring the cost does not exceed the budget. Specifically, λ-generalization method is proposed to lead the policy learning process can be executed according to different λ values adaptively, avoiding re-learning new polices from scratch. Thus the computation overhead is greatly reduced. Further, a novel offline reinforcement learning method and an off-policy evaluation algorithm are proposed for policy learning and policy evaluation, respectively. Finally, experiments on the simulation platform and real-world e-commerce market validate the effectiveness of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80f979c9-e3a8-4950-8b94-4c6f0da43dbbCited by top-tier papers4
- Direct Heterogeneous Causal Learning for Resource Allocation Problems in MarketingHao Zhou, Shaoming Li, Guibin Jiang, Jiaqi Zheng et al.AAAI 2023 · 35 citations
- RL-MPCA: A Reinforcement Learning Based Multi-Phase Computation Allocation Approach for Recommender SystemsJiahong Zhou, Shunhui Mao, Guoliang Yang, Bo Tang et al.WWW 2023 · 10 citations
- Optimizing Marketing Subsidies via Counterfactual Learning with Asymmetric Reward FunctionXiang Li, Yanghao Xiao, Chunyuan Zheng, Qian Zou et al.SIGIR 2026 · 1 citation
- Improve ROI with Causal Learning and Conformal PredictionMeng Ai, Zhuo Chen, Jibin Wang, Jing Shang et al.ICDE 2024
Builds on10
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
- Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision ProcessesDongsheng Ding, Kaiqing Zhang, Tamer Basar, Mihailo R. JovanovicNeurIPS 2020 · 252 citations
- Safe Reinforcement Learning in Constrained Markov Decision ProcessesAkifumi Wachi, Yanan SuiICML 2020 · 190 citations
Related papers
- Off-Policy Learning with Limited SupplyKoichi Tanaka, Ren Kishimoto, Bushun Kawagishi, Yusuke Narita et al.WWW 2026
- SACO: Sequence-Aware Constrained Optimization Framework for Coupon Distribution in E-commerceLi Kong, Bingzhe Wang, Zhou Chen, Suhan Hu et al.AAAI 2026
- Coupon Design in Advertising SystemsWeiran Shen, Pingzhong Tang, Xun Wang, Yadong Xu et al.AAAI 2021 · 5 citations
- Dynamic Knapsack Optimization Towards Efficient Multi-Channel Sequential AdvertisingXiaotian Hao, Zhaoqing Peng, Yi Ma, Guan Wang et al.ICML 2020 · 29 citations
- COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction EstimationJongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz, Nicolas Heess et al.ICLR 2022 · 84 citations
