Sustainable Online Reinforcement Learning for Auto-bidding
Zhiyu Mou, Yusen Huo, Rongquan Bai, Mingzhou Xie, Chuan Yu, Jian Xu, Bo Zheng
摘要
Recently, auto-bidding technique has become an essential tool to increase the revenue of advertisers. Facing the complex and ever-changing bidding environments in the real-world advertising system (RAS), state-of-the-art auto-bidding policies usually leverage reinforcement learning (RL) algorithms to generate real-time bids on behalf of the advertisers. Due to safety concerns, it was believed that the RL training process can only be carried out in an offline virtual advertising system (VAS) that is built based on the historical data generated in the RAS. In this paper, we argue that there exists significant gaps between the VAS and RAS, making the RL training process suffer from the problem of inconsistency between online and offline (IBOO). Firstly, we formally define the IBOO and systematically analyze its causes and influences. Then, to avoid the IBOO, we propose a sustainable online RL (SORL) framework that trains the auto-bidding policy by directly interacting with the RAS, instead of learning in the VAS. Specifically, based on our proof of the Lipschitz smooth property of the Q function, we design a safe and efficient online exploration (SER) policy for continuously collecting data from the RAS. Meanwhile, we derive the theoretical lower bound on the safety of the SER policy. We also develop a variance-suppressed conservative Q-learning (V-CQL) method to effectively and stably learn the auto-bidding policy with the collected data. Finally, extensive simulated and real-world experiments validate the superiority of our approach over the state-of-the-art auto-bidding algorithm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Trajectory-wise Iterative Reinforcement Learning Framework for Auto-biddingHaoming Li, Yusen Huo, Shuai Dou, Zhenzhe Zheng 等WWW 2024 · 被引用 11 次
- Making Offline RL Online: Collaborative World Models for Offline Visual Reinforcement LearningQi Wang, Junming Yang, Yunbo Wang, Xin Jin 等NeurIPS 2024 · 被引用 10 次
- Generative Auto-Bidding with Value-Guided ExplorationsJingtong Gao, Yewen Li, Shuai Mao, Peng Jiang 等SIGIR 2025 · 被引用 7 次
- Enhancing Generative Auto-bidding with Offline Reward Evaluation and Policy SearchZhiyu Mou, Yiqin Lv, Miao Xu, Qi Wang 等ICLR 2026 · 被引用 4 次
- Generative Auto-Bidding with Unified Modeling and ExplorationMingming Zhang, Feiqing Zhuang, Na Li, Shengjie Sun 等SIGIR 2026
它引用的顶会 Paper9
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Critic Regularized RegressionZiyu Wang, Alexander Novikov, Konrad Zolna, Josh Merel 等NeurIPS 2020 · 被引用 406 次
- Offline RL Without Off-Policy EvaluationDavid Brandfonbrener, Will Whitney, Rajesh Ranganath, Joan BrunaNeurIPS 2021 · 被引用 217 次
- BAIL: Best-Action Imitation Learning for Batch Deep Reinforcement LearningXinyue Chen, Zijian Zhou, Zheng Wang, Che Wang 等NeurIPS 2020 · 被引用 146 次
- DeepThermal: Combustion Optimization for Thermal Power Generating Units Using Offline Reinforcement LearningXianyuan Zhan, Haoran Xu, Yue Zhang, Xiangyu Zhu 等AAAI 2022 · 被引用 96 次
相关 Paper
- LBM: Hierarchical Large Auto-Bidding Model via Reasoning and ActingYewen Li, Zhiyi Lyu, Peng Jiang, Qingpeng Cai 等WWW 2026
- Auto-Bidding in Real-Time Auctions via Oracle Imitation LearningAlberto Silvio Chiappa, Briti Gangopadhyay, Zhao Wang, Shingo TakamatsuKDD 2025 · 被引用 4 次
- Auto-bidding under Return-on-Spend Constraints with Uncertainty QuantificationJiale Han, Chun Gan, Chengcheng Zhang, Jie He 等WWW 2026
- DRIVE: Distributional and Retrieval-Augmented Bidding with Value EvaluationMiduo Cui, Haochen Wang, Shangqin Mao, Xun Yang 等ICML 2026 · 被引用 1 次
- On the Coordination of Value-Maximizing BiddersYanru Guan, Jiahao Zhang, Zhe Feng, Tao LinICML 2026
