Cost-Effective Incentive Allocation via Structured Counterfactual Inference
Romain Lopez, Chenchen Li, Xiang Yan, Junwu Xiong, Michael I. Jordan, Yuan Qi, Le Song
Abstract
We address a practical problem ubiquitous in modern marketing campaigns, in which a central agent tries to learn a policy for allocating strategic financial incentives to customers and observes only bandit feedback. In contrast to traditional policy optimization frameworks, we take into account the additional reward structure and budget constraints common in this setting, and develop a new two-step method for solving this constrained counterfactual policy optimization problem. Our method first casts the reward estimation problem as a domain adaptation problem with supplementary structure, and then subsequently uses the estimators for optimizing the policy with constraints. We also establish theoretical error bounds for our estimation procedure and we empirically show that the approach leads to significant improvement on both synthetic and real datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b670a1f1-805f-453c-b33e-3121044e9a55Cited by top-tier papers4
- Learning from eXtreme Bandit FeedbackRomain Lopez, Inderjit S. Dhillon, Michael I. JordanAAAI 2021 · 26 citations
- Off-Policy Evaluation with Policy-Dependent Optimization ResponseWenshuo Guo, Michael I. Jordan, Angela ZhouNeurIPS 2022 · 5 citations
- Enhancing Counterfactual Classification Performance via Self-TrainingRuijiang Gao, Max Biggs, Wei Sun, Ligong HanAAAI 2022 · 3 citations
- Learning Treatment Representations for Downstream Instrumental Variable RegressionShiangyi Lin, Hui Lan, Vasilis SyrgkanisICML 2026
Related papers
- Optimizing Marketing Subsidies via Counterfactual Learning with Asymmetric Reward FunctionXiang Li, Yanghao Xiao, Chunyuan Zheng, Qian Zou et al.SIGIR 2026 · 1 citation
- Causal Inference Under Threshold Manipulation: Bayesian Mixture Modeling and Heterogeneous Treatment EffectsKohsuke Kubota, Shonosuke SugasawaAAAI 2026
- Marketing Hosting: From Fixed to Endogenous BudgetsBingzhe Wang, Tianyu Wang, Qi Qi, Xiaoxuan Deng et al.WWW 2026
- Towards Domain Adaptive Neural Contextual BanditsZiyan Wang, Xiaoming Huo, Hao WangICLR 2025
- A Reduction-Based Framework for Conservative Bandits and Reinforcement LearningYunchang Yang, Tianhao Wu, Han Zhong, Evrard Garcelon et al.ICLR 2022 · 9 citations
