Exploration and Regularization of the Latent Action Space in Recommendation
Shuchang Liu, Qingpeng Cai, Bowen Sun, Yuhao Wang, Ji Jiang, Dong Zheng, Peng Jiang, Kun Gai, Xiangyu Zhao, Yongfeng Zhang
摘要
In recommender systems, reinforcement learning solutions have effectively boosted recommendation performance because of their ability to capture long-term user-system interaction. However, the action space of the recommendation policy is a list of items, which could be extremely large with a dynamic candidate item pool. To overcome this challenge, we propose a hyper-actor and critic learning framework where the policy decomposes the item list generation process into a hyper-action inference step and an effect-action selection step. The first step maps the given state space into a vectorized hyper-action space, and the second step selects the item list based on the hyper-action. In order to regulate the discrepancy between the two action spaces, we design an alignment module along with a kernel mapping function for items to ensure inference accuracy and include a supervision module to stabilize the learning process. We build simulated environments on public datasets and empirically show that our framework is superior in recommendation compared to standard RL baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- LinRec: Linear Attention Mechanism for Long-term Sequential Recommender SystemsLangming Liu, Liu Cai, Chi Zhang, Xiangyu Zhao 等SIGIR 2023 · 被引用 86 次
- Alleviating Matthew Effect of Offline Reinforcement Learning in Interactive RecommendationChongming Gao, Kexin Huang, Jiawei Chen, Yuan Zhang 等SIGIR 2023 · 被引用 65 次
- LLM4Rerank: LLM-based Auto-Reranking Framework for RecommendationsJingtong Gao, Bo Chen, Xiangyu Zhao, Weiwen Liu 等WWW 2025 · 被引用 50 次
- SIGMA: Selective Gated Mamba for Sequential RecommendationZiwei Liu, Qidong Liu, Yejing Wang, Wanyu Wang 等AAAI 2025 · 被引用 31 次
- Explicitly Integrating Judgment Prediction with Legal Document Retrieval: A Law-Guided Generative ApproachWeicong Qin, Zelin Cao, Weijie Yu, Zihua Si 等SIGIR 2024 · 被引用 17 次
它引用的顶会 Paper7
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- Sequential Recommendation with Graph Neural NetworksJianxin Chang, Chen Gao, Yu Zheng, Yiqun Hui 等SIGIR 2021 · 被引用 435 次
- Self-Supervised Reinforcement Learning for Recommender SystemsXin Xin, Alexandros Karatzoglou, Ioannis Arapakis, Joemon M. JoseSIGIR 2020 · 被引用 217 次
- DEAR: Deep Reinforcement Learning for Online Advertising Impression in Recommender SystemsXiangyu Zhao, Changsheng Gu, Haoshenglun Zhang, Xiwang Yang 等AAAI 2021 · 被引用 131 次
- Neural Collaborative ReasoningHanxiong Chen, Shaoyun Shi, Yunqi Li, Yongfeng ZhangWWW 2021 · 被引用 100 次
相关 Paper
- Learning Pseudometric-based Action Representations for Offline Reinforcement LearningPengjie Gu, Mengchen Zhao, Chen Chen, Dong Li 等ICML 2022 · 被引用 17 次
- Unified Conversational Recommendation Policy Learning via Graph-based Reinforcement LearningYang Deng, Yaliang Li, Fei Sun, Bolin Ding 等SIGIR 2021 · 被引用 131 次
- Reinforcement Learning-Constrained Segmented User Modeling with Large Language Models for RecommendationYu Xia, Qing Tan, He Chen, Jingyu Chen 等WWW 2026
- Aligning Large Language Models for Controllable RecommendationsWensheng Lu, Jianxun Lian, Wei Zhang, Guanghua Li 等ACL 2024 · 被引用 7 次
- Reinforcement Learning with a Disentangled Universal Value Function for Item RecommendationKai Wang, Zhene Zou, Qilin Deng, Jianrong Tao 等AAAI 2021 · 被引用 25 次
