Learning Personalized Ad Impact via Contextual Reinforcement Learning under Delayed Rewards
Yuwei Cheng, Zifeng Zhao, Haifeng Xu
摘要
Online advertising platforms use automated auctions to connect advertisers with potential customers, requiring effective bidding strategies to maximize profits. Accurate ad impact estimation requires considering three key factors: delayed and long-term effects, cumulative ad impacts such as reinforcement or fatigue, and customer heterogeneity. However, these effects are often not jointly addressed in previous studies. To capture these factors, we model ad bidding as a Contextual Markov Decision Process (CMDP) with delayed Poisson rewards. For efficient estimation, we propose a two-stage maximum likelihood estimator combined with data-splitting strategies, ensuring controlled estimation error based on the first-stage estimator's (in)accuracy. Building on this, we design a reinforcement learning algorithm to derive efficient personalized bidding strategies. This approach achieves a near-optimal regret bound of , where is the contextual dimension, is the number of rounds, and is the number of customers. Our theoretical findings are validated by simulation experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Nearly Optimal Algorithms for Linear Contextual Bandits with Adversarial CorruptionsJiafan He, Dongruo Zhou, Tong Zhang, Quanquan GuNeurIPS 2022 · 被引用 66 次
- Learning to Bid in Contextual First Price Auctions✱Ashwinkumar Badanidiyuru, Zhe Feng, Guru GuruganeshWWW 2023 · 被引用 24 次
- Learning Neural Contextual Bandits through Perturbed RewardsYiling Jia, Weitong Zhang, Dongruo Zhou, Quanquan Gu 等ICLR 2022 · 被引用 20 次
- Efficient Algorithms for Generalized Linear Bandits with Heavy-tailed RewardsBo Xue, Yimu Wang, Yuanyu Wan, Jinfeng Yi 等NeurIPS 2023 · 被引用 16 次
- Online Second Price Auction with Semi-Bandit Feedback under the Non-Stationary SettingHaoyu Zhao, Wei ChenAAAI 2020 · 被引用 15 次
相关 Paper
- User Response in Ad Auctions: An MDP Formulation of Long-term Revenue OptimizationYang Cai, Zhe Feng, Christopher Liaw, Aranyak Mehta 等WWW 2024 · 被引用 1 次
- Incrementality Bidding via Reinforcement Learning under Mixed and Delayed RewardsAshwinkumar Badanidiyuru Varadaraja, Zhe Feng, Tianxi Li, Haifeng XuNeurIPS 2022 · 被引用 7 次
- No-Regret Online Autobidding Algorithms in First-price AuctionsYilin Li, Yuan Deng, Wei Tang, Hanrui ZhangNeurIPS 2025 · 被引用 4 次
- No-Regret Algorithms in non-Truthful Auctions with Budget and ROI ConstraintsGagan Aggarwal, Giannis Fikioris, Mingfei ZhaoWWW 2025 · 被引用 13 次
- Online Ad Procurement in Non-stationary Autobidding WorldsJason Cheuk Nam Liang, Haihao Lu, Baoyu ZhouNeurIPS 2023 · 被引用 10 次
