MaHRL: Multi-goals Abstraction Based Deep Hierarchical Reinforcement Learning for Recommendations
Dongyang Zhao, Liang Zhang, Bo Zhang, Lizhou Zheng, Yongjun Bao, Weipeng Yan
摘要
As huge commercial value of the recommender system, there has been growing interest to improve its performance in recent years. The majority of existing methods have achieved great improvement on the metric of click, but perform poorly on the metric of conversion possibly due to its extremely sparse feedback signal. To track this challenge, we design a novel deep hierarchical reinforcement learning based recommendation framework to model consumers' hierarchical purchase interest. Specifically, the high-level agent catches long-term sparse conversion interest, and automatically sets abstract goals for low-level agent, while the low-level agent follows the abstract goals and catches short-term click interest via interacting with real-time environment. To solve the inherent problem in hierarchical reinforcement learning, we propose a novel multi-goals abstraction based deep hierarchical reinforcement learning algorithm (MaHRL). Our proposed algorithm contains three contributions: 1) the high-level agent generates multiple goals to guide the low-level agent in different sub-periods, which reduces the difficulty of approaching high-level goals; 2) different goals share the same state encoder structure and its parameters, which increases the update frequency of the high-level agent and thus accelerates the convergence of our proposed algorithm; 3) an appreciated reward assignment mechanism is designed to allocate rewards in each goal so as to coordinate different goals in a consistent direction. We evaluate our proposed algorithm based on a real-world e-commerce dataset and validate its effectiveness.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Hierarchical Reinforcement Learning for Integrated RecommendationRuobing Xie, Shaoliang Zhang, Rui Wang, Feng Xia 等AAAI 2021 · 被引用 90 次
- PrefRec: Recommender Systems with Human Preferences for Reinforcing Long-term User EngagementWanqi Xue, Qingpeng Cai, Zhenghai Xue, Shuo Sun 等KDD 2023 · 被引用 24 次
- ResAct: Reinforcing Long-term Engagement in Sequential Recommendation with Residual ActorWanqi Xue, Qingpeng Cai, Ruohan Zhan, Dong Zheng 等ICLR 2023 · 被引用 6 次
- DARLR: Dual-Agent Offline Reinforcement Learning for Recommender Systems with Dynamic RewardYi Zhang, Ruihong Qiu, Xuwei Xu, Jiajun Liu 等SIGIR 2025 · 被引用 3 次
相关 Paper
- Confident Action Decision via Hierarchical Policy Learning for Conversational RecommendationHeeseon Kim, Hyeongjun Yang, Kyong-Ho LeeWWW 2023 · 被引用 10 次
- Reinforcement Learning with a Disentangled Universal Value Function for Item RecommendationKai Wang, Zhene Zou, Qilin Deng, Jianrong Tao 等AAAI 2021 · 被引用 25 次
- DEAR: Deep Reinforcement Learning for Online Advertising Impression in Recommender SystemsXiangyu Zhao, Changsheng Gu, Haoshenglun Zhang, Xiwang Yang 等AAAI 2021 · 被引用 131 次
- HutCRS: Hierarchical User-Interest Tracking for Conversational Recommender SystemMingjie Qian, Yongsen Zheng, Jinghui Qin, Liang LinEMNLP 2023 · 被引用 11 次
- Hierarchical Reinforcement Learning with Targeted Causal InterventionsMohammadsadegh Khorasani, Saber Salehkaleybar, Negar Kiyavash, Matthias GrossglauserICML 2025
