UNEX-RL: Reinforcing Long-Term Rewards in Multi-Stage Recommender Systems with UNidirectional EXecution
Gengrui Zhang, Yao Wang, Xiaoshuang Chen, Hongyi Qian, Kaiqiao Zhan, Ben Wang
摘要
In recent years, there has been a growing interest in utilizing reinforcement learning (RL) to optimize long-term rewards in recommender systems. Since industrial recommender systems are typically designed as multi-stage systems, RL methods with a single agent face challenges when optimizing multiple stages simultaneously. The reason is that different stages have different observation spaces, and thus cannot be modeled by a single agent. To address this issue, we propose a novel UNidirectional-EXecution-based multi-agent Reinforcement Learning (UNEX-RL) framework to reinforce the long-term rewards in multi-stage recommender systems. We show that the unidirectional execution is a key feature of multi-stage recommender systems, bringing new challenges to the applications of multi-agent reinforcement learning (MARL), namely the observation dependency and the cascading effect. To tackle these challenges, we provide a cascading information chain (CIC) method to separate the independent observations from action-dependent observations and use CIC to train UNEX-RL effectively. We also discuss practical variance reduction techniques for UNEX-RL. Finally, we show the effectiveness of UNEX-RL on both public datasets and an online recommender system with over 100 million users. Specifically, UNEX-RL reveals a 0.558% increase in users' usage time compared with single-agent RL algorithms in online A/B experiments, highlighting the effectiveness of UNEX-RL in industrial recommender systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank SystemsRuoxi Wang, Rakesh Shivanna, Derek Zhiyuan Cheng, Sagar Jain 等WWW 2021 · 被引用 793 次
- FACMAC: Factored Multi-Agent Centralised Policy GradientsBei Peng, Tabish Rashid, Christian Schröder de Witt, Pierre-Alexandre Kamienny 等NeurIPS 2021 · 被引用 399 次
- DOP: Off-Policy Multi-Agent Decomposed Policy GradientsYihan Wang, Beining Han, Tonghan Wang, Heng Dong 等ICLR 2021 · 被引用 208 次
- PrefRec: Recommender Systems with Human Preferences for Reinforcing Long-term User EngagementWanqi Xue, Qingpeng Cai, Zhenghai Xue, Shuo Sun 等KDD 2023 · 被引用 24 次
- ResAct: Reinforcing Long-term Engagement in Sequential Recommendation with Residual ActorWanqi Xue, Qingpeng Cai, Ruohan Zhan, Dong Zheng 等ICLR 2023 · 被引用 6 次
相关 Paper
- Can Cooperative Multi-Agent Reinforcement Learning Boost Automatic Web Testing? An Exploratory StudyYujia Fan, Sinan Wang, Zebang Fei, Yao Qin 等ASE 2024 · 被引用 3 次
- MARLIN: Multi-Agent Reinforcement Learning for Incremental DAG DiscoveryDong Li, Zhengzhang Chen, Xujiang Zhao, Linlin Yu 等AAAI 2026
- MaHRL: Multi-goals Abstraction Based Deep Hierarchical Reinforcement Learning for RecommendationsDongyang Zhao, Liang Zhang, Bo Zhang, Lizhou Zheng 等SIGIR 2020 · 被引用 33 次
- Reinforcement Learning with a Disentangled Universal Value Function for Item RecommendationKai Wang, Zhene Zou, Qilin Deng, Jianrong Tao 等AAAI 2021 · 被引用 25 次
- RecFlow: An Industrial Full Flow Recommendation DatasetQi Liu, Kai Zheng, Rui Huang, Wuchao Li 等ICLR 2025
