VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement Learning
Siran Chen, Boyu Chen, Yuxiao Luo, Chenyun Yu, Yi Ouyang, Lei Cheng, Chengxiang Zhuo, Zang Li, Yali Wang
摘要
Owing to powerful natural language processing and generative capabilities, large language model (LLM) agents have emerged as a promising solution for enhancing recommendation systems via user simulation. However, in the realm of video recommendation, existing studies predominantly resort to prompt-based simulation using frozen LLMs and encounter the intricate challenge of multimodal content understanding. This frequently results in suboptimal item modeling and user preference learning, thereby ultimately constraining recommendation performance. To address these challenges, we introduce VRAgent-R1, a novel agent-based paradigm that incorporates human-like intelligence in user simulation. Specifically, VRAgent-R1 comprises two distinct agents: the Item Perception (IP) Agent and the User Simulation (US) Agent, designed for interactive user-item modeling. Firstly, the IP Agent emulates human-like progressive thinking based on MLLMs, effectively capturing hidden recommendation semantics in videos. With a more comprehensive multimodal content understanding provided by the IP Agent, the video recommendation system is equipped to provide higher-quality candidate items. Subsequently, the US Agent refines the recommended video sets based on in-depth chain-of-thought (CoT) reasoning and achieves better alignment with real user preferences through reinforcement learning. Experimental results on a large-scale video recommendation benchmark have demonstrated the effectiveness of our proposed VRAgent-R1 method, e.g., the IP Agent achieves a 6.0% improvement in NDCG@10 on the MicroLens-100k dataset, while the US Agent shows approximately 45.0% higher accuracy in user decision simulation compared to state-of-the-art baselines.
Preprint. Under review. This is a video titled "in the middle of the night my pig addiction again committed # filial police Art # food # drama", will the user like the video? MLLM: a man wearing a black hoodie with yellow letter…, maybe related to Chinese content or culture. SFT: <answer> No </answer>.
IP Agent: a humorous and exaggerated drama shows three men in black are seeking and arresting people with pig addiction .... #fictional drama #humorous This is a video titled "The situation suddenly changed", will the user like the video? SFT Ours US Agent: <think> based on the use's historic…, shows a positive attitude…may like pleasant, relaxing, humorous content </think> <answer> Yes </answer>.
Ours US Agent: <think> based on the use's historic…, the user loves games, sports and TV shows, not shows interest in politic topics </think> <answer> No </answer>.
MLLM: two men wearing suit, the old man in the center is smiling and seems friendly… SFT: <answer> Yes </answer>.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and GenerationZhengrong Yue, Haiyu Zhang, Xiangyu Zeng, Boyu Chen 等ICLR 2026 · 被引用 25 次
- LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM AgentsBoyu Chen, Zhengrong Yue, Siran Chen, Zikang Wang 等ICCV 2025 · 被引用 12 次
- VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement LearningBoyu Chen, Zikang Wang, Zhengrong Yue, Kainan Yan 等CVPR 2026 · 被引用 11 次
- G-UBS: Towards Robust Understanding of Implicit Feedback via Group-Aware User Behavior SimulationBoyu Chen, Siran Chen, Zhengrong Yue, Kainan Yan 等AAAI 2026 · 被引用 7 次
- When Top-ranked Recommendations Fail: Modeling Multi-Granular Negative Feedback for Explainable and Robust Video RecommendationSiran Chen, Boyu Chen, Chenyun Yu, Yi Ouyang 等AAAI 2026 · 被引用 5 次
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li 等SIGIR 2020 · 被引用 4,448 次
- Visual-RFT: Visual Reinforcement Fine-TuningZiyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong 等ICCV 2025 · 被引用 563 次
- Representation Learning with Large Language Models for RecommendationXubin Ren, Wei Wei, Lianghao Xia, Lixin Su 等WWW 2024 · 被引用 385 次
- Parameter-Efficient Transfer from Sequential Behaviors for User Modeling and RecommendationFajie Yuan, Xiangnan He, Alexandros Karatzoglou, Liguang ZhangSIGIR 2020 · 被引用 155 次
相关 Paper
- LLM-Powered User Simulator for Recommender SystemZijian Zhang, Shuchang Liu, Ziru Liu, Rui Zhong 等AAAI 2025 · 被引用 11 次
- Agentic Feedback Loop Modeling Improves Recommendation and User SimulationShihao Cai, Jizhi Zhang, Keqin Bao, Chongming Gao 等SIGIR 2025 · 被引用 13 次
- MLLMRec: A Preference Reasoning Paradigm with Graph Refinement for Multimodal RecommendationYuzhuo Dang, Xin Zhang, Zhiqiang Pan, Yuxiao Duan 等SIGIR 2026 · 被引用 1 次
- Task-Aware Automated User Profile Generation for Recommendation Simulation Using Large Language ModelsXinye Wanyan, Chenglong Ma, Danula Hettiachchi, Ziqi Xu 等SIGIR 2026
- Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video UnderstandingXiaoqian Shen, Wenxuan Zhang, Jun Chen, Mohamed ElhoseinyNeurIPS 2025 · 被引用 37 次
