Optimizing Multi-Turn Interactive Recommendation Agents via Generative Intrinsic Motivation
Xueyang Feng, Jiakai Tang, Xu Chen, Quanyu Dai, Zhenhua Dong
Abstract
Large language models have given rise to interactive recommendation agents (IRAs). Through proactive clarification, tool invocation, and dynamic dialogue, IRAs shift recommender systems from passive prediction to interactive, proactive intelligence. For training IRAs, agentic reinforcement learning offers a natural pathway, as it enables models to learn interactive capabilities directly from environmental feedback without requiring costly annotated data. However, this process faces three key challenges: credit assignment in multi-turn interactions, efficient exploration in large action spaces, and coordinated learning of multiple interactive skills.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Action First: Leveraging Preference-Aware Actions for More Effective Decision-Making in Interactive Recommender SystemsRenting Rui, Yunjia Xi, Weiwen Liu, Jianghao Lin et al.SIGIR 2025 · 1 citation
- Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic OptimizationYihang Yao, Zhepeng Cen, Haohong Lin, Shiqi Liu et al.ICML 2026
- Agentic Reinforcement Learning with Implicit Step RewardsXiaoqian Liu, Ke Wang, Yuchuan Wu, Fei Huang et al.ICLR 2026 · 46 citations
- AgentGym-RL: An Open-Source Framework to Train LLM Agents for Long-Horizon Decision Making via Multi-Turn RLZhiheng Xi, Jixuan Huang, Chenyang Liao, Baodai Huang et al.ICLR 2026
- Aligning Large Language Models for Controllable RecommendationsWensheng Lu, Jianxun Lian, Wei Zhang, Guanghua Li et al.ACL 2024 · 7 citations
