Adversarial Gradient Driven Exploration for Deep Click-Through Rate Prediction
Kailun Wu, Weijie Bian, Zhangming Chan, Lejian Ren, Shiming Xiang, Shuguang Han, Hongbo Deng, Bo Zheng
Abstract
Exploration-Exploitation (E& E) algorithms are commonly adopted to deal with the feedback-loop issue in large-scale online recommender systems. Most of existing studies believe that high uncertainty can be a good indicator of potential reward, and thus primarily focus on the estimation of model uncertainty. We argue that such an approach overlooks the subsequent effect of exploration on model training. From the perspective of online learning, the adoption of an exploration strategy would also affect the collecting of training data, which further influences model learning. To understand the interaction between exploration and training, we design a Pseudo-Exploration module that simulates the model updating process after a certain item is explored and the corresponding feedback is received. We further show that such a process is equivalent to adding an adversarial perturbation to the model input, and thereby name our proposed approach as an the Adversarial Gradient Driven Exploration (AGE). For production deployment, we propose a dynamic gating unit to pre-determine the utility of an exploration. This enables us to utilize the limited amount of resources for exploration, and avoid wasting pageview resources on ineffective exploration. The effectiveness of AGE was firstly examined through an extensive number of ablation studies on an academic dataset. Meanwhile, AGE has also been deployed to one of the world-leading display advertising platforms, and we observe significant improvements on various top-line evaluation metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7f7d3b3d-6027-4e59-b106-f8dc19103d16Cited by top-tier papers2
- Reducing Symbiosis Bias through Better A/B Tests of Recommendation AlgorithmsJennifer Brennan, Yahu Cong, Yiwei Yu, Lina Lin et al.WWW 2025 · 8 citations
- Bayesian Ensemble for Sequential Decision-MakingRui Liu, Enmin Zhao, Lu Wang, Yu Li et al.ICLR 2026 · 2 citations
Builds on7
- Neural Contextual Bandits with UCB-based ExplorationDongruo Zhou, Lihong Li, Quanquan GuICML 2020 · 329 citations
- Neural Thompson SamplingWeitong Zhang, Dongruo Zhou, Lihong Li, Quanquan GuICLR 2021 · 152 citations
- How to Retrain Recommender System?: A Sequential Meta-Learning MethodYang Zhang, Fuli Feng, Chenxu Wang, Xiangnan He et al.SIGIR 2020 · 70 citations
- UKD: Debiasing Conversion Rate Estimation via Uncertainty-regularized Knowledge DistillationZixuan Xu, Penghui Wei, Weimin Zhang, Shaoguo Liu et al.WWW 2022 · 31 citations
- Selection and Generation: Learning towards Multi-Product Advertisement Post GenerationZhangming Chan, Yuchi Zhang, Xiuying Chen, Shen Gao et al.EMNLP 2020 · 26 citations
Related papers
- Trajectory-wise Iterative Reinforcement Learning Framework for Auto-biddingHaoming Li, Yusen Huo, Shuai Dou, Zhenzhe Zheng et al.WWW 2024 · 11 citations
- Generative Explore-Exploit: Training-free Optimization of Generative Recommender Systems using LLM OptimizersLütfi Kerem Senel, Besnik Fetahu, Davis Yoshida, Zhiyu Chen et al.ACL 2024
- An Adversarial Imitation Click Model for Information RetrievalXinyi Dai, Jianghao Lin, Weinan Zhang, Shuai Li et al.WWW 2021 · 40 citations
- DEAR: Deep Reinforcement Learning for Online Advertising Impression in Recommender SystemsXiangyu Zhao, Changsheng Gu, Haoshenglun Zhang, Xiwang Yang et al.AAAI 2021 · 131 citations
- GuideBoot: Guided Bootstrap for Deep Contextual Banditsin Online AdvertisingFeiyang Pan, Haoming Li, Xiang Ao, Wei Wang et al.WWW 2021
