VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement Learning
Siran Chen, Boyu Chen, Yuxiao Luo, Chenyun Yu, Yi Ouyang, Lei Cheng, Chengxiang Zhuo, Zang Li, Yali Wang
Abstract
Owing to powerful natural language processing and generative capabilities, large language model (LLM) agents have emerged as a promising solution for enhancing recommendation systems via user simulation. However, in the realm of video recommendation, existing studies predominantly resort to prompt-based simulation using frozen LLMs and encounter the intricate challenge of multimodal content understanding. This frequently results in suboptimal item modeling and user preference learning, thereby ultimately constraining recommendation performance. To address these challenges, we introduce VRAgent-R1, a novel agent-based paradigm that incorporates human-like intelligence in user simulation. Specifically, VRAgent-R1 comprises two distinct agents: the Item Perception (IP) Agent and the User Simulation (US) Agent, designed for interactive user-item modeling. Firstly, the IP Agent emulates human-like progressive thinking based on MLLMs, effectively capturing hidden recommendation semantics in videos. With a more comprehensive multimodal content understanding provided by the IP Agent, the video recommendation system is equipped to provide higher-quality candidate items. Subsequently, the US Agent refines the recommended video sets based on in-depth chain-of-thought (CoT) reasoning and achieves better alignment with real user preferences through reinforcement learning. Experimental results on a large-scale video recommendation benchmark have demonstrated the effectiveness of our proposed VRAgent-R1 method, e.g., the IP Agent achieves a 6.0% improvement in NDCG@10 on the MicroLens-100k dataset, while the US Agent shows approximately 45.0% higher accuracy in user decision simulation compared to state-of-the-art baselines.
Preprint. Under review. This is a video titled "in the middle of the night my pig addiction again committed # filial police Art # food # drama", will the user like the video? MLLM: a man wearing a black hoodie with yellow letter…, maybe related to Chinese content or culture. SFT: <answer> No </answer>.
IP Agent: a humorous and exaggerated drama shows three men in black are seeking and arresting people with pig addiction .... #fictional drama #humorous This is a video titled "The situation suddenly changed", will the user like the video? SFT Ours US Agent: <think> based on the use's historic…, shows a positive attitude…may like pleasant, relaxing, humorous content </think> <answer> Yes </answer>.
Ours US Agent: <think> based on the use's historic…, the user loves games, sports and TV shows, not shows interest in politic topics </think> <answer> No </answer>.
MLLM: two men wearing suit, the old man in the center is smiling and seems friendly… SFT: <answer> Yes </answer>.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d6eca36-0fc1-49f1-83b0-ae48d6b5447aCited by top-tier papers6
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and GenerationZhengrong Yue, Haiyu Zhang, Xiangyu Zeng, Boyu Chen et al.ICLR 2026 · 25 citations
- LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM AgentsBoyu Chen, Zhengrong Yue, Siran Chen, Zikang Wang et al.ICCV 2025 · 12 citations
- VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement LearningBoyu Chen, Zikang Wang, Zhengrong Yue, Kainan Yan et al.CVPR 2026 · 11 citations
- G-UBS: Towards Robust Understanding of Implicit Feedback via Group-Aware User Behavior SimulationBoyu Chen, Siran Chen, Zhengrong Yue, Kainan Yan et al.AAAI 2026 · 7 citations
- When Top-ranked Recommendations Fail: Modeling Multi-Granular Negative Feedback for Explainable and Robust Video RecommendationSiran Chen, Boyu Chen, Chenyun Yu, Yi Ouyang et al.AAAI 2026 · 5 citations
Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- Visual-RFT: Visual Reinforcement Fine-TuningZiyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong et al.ICCV 2025 · 563 citations
- Representation Learning with Large Language Models for RecommendationXubin Ren, Wei Wei, Lianghao Xia, Lixin Su et al.WWW 2024 · 385 citations
- Parameter-Efficient Transfer from Sequential Behaviors for User Modeling and RecommendationFajie Yuan, Xiangnan He, Alexandros Karatzoglou, Liguang ZhangSIGIR 2020 · 155 citations
Related papers
- LLM-Powered User Simulator for Recommender SystemZijian Zhang, Shuchang Liu, Ziru Liu, Rui Zhong et al.AAAI 2025 · 11 citations
- Agentic Feedback Loop Modeling Improves Recommendation and User SimulationShihao Cai, Jizhi Zhang, Keqin Bao, Chongming Gao et al.SIGIR 2025 · 13 citations
- MLLMRec: A Preference Reasoning Paradigm with Graph Refinement for Multimodal RecommendationYuzhuo Dang, Xin Zhang, Zhiqiang Pan, Yuxiao Duan et al.SIGIR 2026 · 1 citation
- Task-Aware Automated User Profile Generation for Recommendation Simulation Using Large Language ModelsXinye Wanyan, Chenglong Ma, Danula Hettiachchi, Ziqi Xu et al.SIGIR 2026
- Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video UnderstandingXiaoqian Shen, Wenxuan Zhang, Jun Chen, Mohamed ElhoseinyNeurIPS 2025 · 37 citations
