Generative Auto-Bidding with Value-Guided Explorations
Jingtong Gao, Yewen Li, Shuai Mao, Peng Jiang, Nan Jiang, Yejing Wang, Qingpeng Cai, Fei Pan, Peng Jiang, Kun Gai, Bo An, Xiangyu Zhao
Abstract
Auto-bidding, with its strong capability to optimize bidding decisions within dynamic and competitive online environments, has become a pivotal strategy for advertising platforms. Existing approaches typically employ rule-based strategies or Reinforcement Learning (RL) techniques. However, rule-based strategies lack the flexibility to adapt to time-varying market conditions, and RL-based methods struggle to capture essential historical dependencies and observations within Markov Decision Process (MDP) frameworks. Furthermore, these approaches often face challenges in ensuring strategy adaptability across diverse advertising objectives. Additionally, as offline training methods are increasingly adopted to facilitate the deployment and maintenance of stable online strategies, the issues of documented behavioral patterns and behavioral collapse resulting from training on fixed offline datasets become increasingly significant. To address these limitations, this paper introduces a novel offline Generative Auto-bidding framework with Value-Guided Explorations (GAVE). GAVE accommodates various advertising objectives through a score-based Return-To-Go (RTG) module. Moreover, GAVE integrates an action exploration mechanism with an RTG-based evaluation method to explore novel actions while ensuring stability-preserving updates. A learnable value function is also designed to guide the direction of action exploration and mitigate Out-of-Distribution (OOD) problems. Experimental results on two offline datasets and real-world deployments demonstrate that GAVE outperforms state-of-the-art baselines in both offline evaluations and online A/B tests. By applying the core methods of this framework, we proudly secured first place in the NeurIPS 2024 competition, 'AIGB Track: Learning Auto-Bidding Agents with Generative Models'.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4c1aa2e5-cb81-4e7b-9063-e51a4203c3b4Cited by top-tier papers7
- GFlowGR: Fine-tuning Generative Recommendation Frameworks with Generative Flow NetworksYejing Wang, Shengyu Zhou, Jinyu Lu, Qidong Liu et al.SIGIR 2026 · 1 citation
- DRIVE: Distributional and Retrieval-Augmented Bidding with Value EvaluationMiduo Cui, Haochen Wang, Shangqin Mao, Xun Yang et al.ICML 2026 · 1 citation
- Generative Auto-Bidding with Unified Modeling and ExplorationMingming Zhang, Feiqing Zhuang, Na Li, Shengjie Sun et al.SIGIR 2026
- AHBid: An Adaptable Hierarchical Bidding Framework for Cross-Channel AdvertisingXinxin Yang, Yangyang Tang, Yikun Zhou, Yaolei Liu et al.WWW 2026
- Hierarchical Residual Policy Optimization for Generative RecommendationsKaifeng Guo, Yiming Yang, Jingtong Gao, Guolei Zeng et al.KDD 2026
Builds on18
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- Offline Reinforcement Learning with Fisher Divergence Critic RegularizationIlya Kostrikov, Rob Fergus, Jonathan Tompson, Ofir NachumICML 2021 · 350 citations
- Online Decision TransformerQinqing Zheng, Amy Zhang, Aditya GroverICML 2022 · 256 citations
Related papers
- Enhancing Generative Auto-bidding with Offline Reward Evaluation and Policy SearchZhiyu Mou, Yiqin Lv, Miao Xu, Qi Wang et al.ICLR 2026 · 4 citations
- LBM: Hierarchical Large Auto-Bidding Model via Reasoning and ActingYewen Li, Zhiyi Lyu, Peng Jiang, Qingpeng Cai et al.WWW 2026
- Sustainable Online Reinforcement Learning for Auto-biddingZhiyu Mou, Yusen Huo, Rongquan Bai, Mingzhou Xie et al.NeurIPS 2022 · 53 citations
- Trajectory-wise Iterative Reinforcement Learning Framework for Auto-biddingHaoming Li, Yusen Huo, Shuai Dou, Zhenzhe Zheng et al.WWW 2024 · 11 citations
- On the Coordination of Value-Maximizing BiddersYanru Guan, Jiahao Zhang, Zhe Feng, Tao LinICML 2026
