Large-Scale Online Learning for Generative List Recommendation in E-commerce: An Environment Policy Optimization Approach
Yuan Wang, Zhiyu Li, Ang Gao, Changshuo Zhang, Xiao Zhang, Jun Xu, Quan Lin
摘要
Generative List Recommendation (GLR) models have shown superior performance in E-commerce by directly generating high-quality recommendation lists through sequential item selection. While Online Learning (OL) has proven valuable for streaming point-wise recommendation models in adapting to dynamic user preferences, its application to GLR remains largely unexplored, due in large part to the inefficiency and instability of conventional on-policy reinforcement learning algorithms used in existing GLR approaches. Existing approaches typically rely on surrogate losses, which provide indirect and biased gradient estimates, making them ill-suited for the rapid, subtle distribution shifts common in real-world E-commerce environments. In this paper, we propose Environment Policy Optimization (EPO), a novel GLR model that fundamentally reshapes policy learning by exploiting the differentiability of the environment within the Generator-Evaluator framework. EPO recognizes that the evaluator is a neural network, capable of providing gradient signals. By directly utilizing these gradients, EPO enables end-to-end optimization of the total list-wise reward—the true objective. To ensure differentiable list generation, EPO introduces a new indexing and exploration strategy based on NeuralSort and Gumbel noise, which relaxes discrete item selection into a continuous, gradient-friendly operation. EPO not only demonstrates strong performance in offline evaluations but also unlocks the potential of online learning for GLR at an industrial scale, yielding a 1.18% relative improvement in user clicks in online A/B tests. The results underscore the critical role of EPO in providing the sensitivity, stability, and gradient fidelity necessary for effective real-time adaptation in streaming and dynamic recommendation environments. Moreover, EPO reached the baseline performance with a training time deduction of 76% under identical hardware conditions.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Differentiable Semantic ID for Generative RecommendationJunchen Fu, Xuri Ge, Alexandros Karatzoglou, Ioannis Arapakis 等SIGIR 2026
- Generative Flow Network for Listwise RecommendationShuchang Liu, Qingpeng Cai, Zhankui He, Bowen Sun 等KDD 2023 · 被引用 15 次
- PiRank: Scalable Learning To Rank via Differentiable SortingRobin M. E. Swezey, Aditya Grover, Bruno Charron, Stefano ErmonNeurIPS 2021 · 被引用 45 次
- Locality-Sensitive State-Guided Experience Replay Optimization for Sparse Rewards in Online RecommendationXiaocong Chen, Lina Yao, Julian J. McAuley, Weili Guan 等SIGIR 2022 · 被引用 14 次
- Graph Diffusion Policy OptimizationYijing Liu, Chao Du, Tianyu Pang, Chongxuan Li 等NeurIPS 2024 · 被引用 23 次
