Exploration and Regularization of the Latent Action Space in Recommendation
Shuchang Liu, Qingpeng Cai, Bowen Sun, Yuhao Wang, Ji Jiang, Dong Zheng, Peng Jiang, Kun Gai, Xiangyu Zhao, Yongfeng Zhang
Abstract
In recommender systems, reinforcement learning solutions have effectively boosted recommendation performance because of their ability to capture long-term user-system interaction. However, the action space of the recommendation policy is a list of items, which could be extremely large with a dynamic candidate item pool. To overcome this challenge, we propose a hyper-actor and critic learning framework where the policy decomposes the item list generation process into a hyper-action inference step and an effect-action selection step. The first step maps the given state space into a vectorized hyper-action space, and the second step selects the item list based on the hyper-action. In order to regulate the discrepancy between the two action spaces, we design an alignment module along with a kernel mapping function for items to ensure inference accuracy and include a supervision module to stabilize the learning process. We build simulated environments on public datasets and empirically show that our framework is superior in recommendation compared to standard RL baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7f52ca6c-ac44-4d69-be13-f14fc09271efCited by top-tier papers17
- LinRec: Linear Attention Mechanism for Long-term Sequential Recommender SystemsLangming Liu, Liu Cai, Chi Zhang, Xiangyu Zhao et al.SIGIR 2023 · 86 citations
- Alleviating Matthew Effect of Offline Reinforcement Learning in Interactive RecommendationChongming Gao, Kexin Huang, Jiawei Chen, Yuan Zhang et al.SIGIR 2023 · 65 citations
- LLM4Rerank: LLM-based Auto-Reranking Framework for RecommendationsJingtong Gao, Bo Chen, Xiangyu Zhao, Weiwen Liu et al.WWW 2025 · 50 citations
- SIGMA: Selective Gated Mamba for Sequential RecommendationZiwei Liu, Qidong Liu, Yejing Wang, Wanyu Wang et al.AAAI 2025 · 31 citations
- Explicitly Integrating Judgment Prediction with Legal Document Retrieval: A Law-Guided Generative ApproachWeicong Qin, Zelin Cao, Weijie Yu, Zihua Si et al.SIGIR 2024 · 17 citations
Builds on7
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- Sequential Recommendation with Graph Neural NetworksJianxin Chang, Chen Gao, Yu Zheng, Yiqun Hui et al.SIGIR 2021 · 435 citations
- Self-Supervised Reinforcement Learning for Recommender SystemsXin Xin, Alexandros Karatzoglou, Ioannis Arapakis, Joemon M. JoseSIGIR 2020 · 217 citations
- DEAR: Deep Reinforcement Learning for Online Advertising Impression in Recommender SystemsXiangyu Zhao, Changsheng Gu, Haoshenglun Zhang, Xiwang Yang et al.AAAI 2021 · 131 citations
- Neural Collaborative ReasoningHanxiong Chen, Shaoyun Shi, Yunqi Li, Yongfeng ZhangWWW 2021 · 100 citations
Related papers
- Learning Pseudometric-based Action Representations for Offline Reinforcement LearningPengjie Gu, Mengchen Zhao, Chen Chen, Dong Li et al.ICML 2022 · 17 citations
- Unified Conversational Recommendation Policy Learning via Graph-based Reinforcement LearningYang Deng, Yaliang Li, Fei Sun, Bolin Ding et al.SIGIR 2021 · 131 citations
- Reinforcement Learning-Constrained Segmented User Modeling with Large Language Models for RecommendationYu Xia, Qing Tan, He Chen, Jingyu Chen et al.WWW 2026
- Aligning Large Language Models for Controllable RecommendationsWensheng Lu, Jianxun Lian, Wei Zhang, Guanghua Li et al.ACL 2024 · 7 citations
- Reinforcement Learning with a Disentangled Universal Value Function for Item RecommendationKai Wang, Zhene Zou, Qilin Deng, Jianrong Tao et al.AAAI 2021 · 25 citations
