Reality vs Counterfactual: Multi-World Contrastive Reinforcement Learning for Enhancing MLLM's Theory of Mind in Egocentric Videos
Guiyang Hou, Yihui Fu, Chen Wu, Xiang Huang, Zhe Zheng, Wenqi Zhang, Yongliang Shen, Weiming Lu
摘要
Theory of Mind (ToM) refers to the ability to infer others' mental states, which is an essential capability for embodied AI agents to effectively collaborate and interact with humans. While improving Large Language Models' ability to reason about characters' mental states in text-based stories/dialogues has been extensively studied, enhancing Multimodal Large Language Models' ToM capabilities, particularly in egocentric video from an embodied perspective, remains unexplored. In this paper, we propose a contrastive Reinforcement Learning (RL) paradigm that explicitly encourages models to leverage temporal and causal evolutionary patterns in user action sequences to infer user's mental states (goals, beliefs, and potential next actions). Evaluation results on in-domain and out-of-domain demonstrate that our method achieves performance improvements of (+30.00%, +2.00%) and (+5.83%, +5.00%) compared to the backbone model and vanilla Group Relative Policy Optimization (GRPO) model, respectively. Additionally, we compare the performance of two post-training paradigms (Supervise Fine-Tuning and RL) and systematically analyze the reasoning trajectories across the base model, vanilla GRPO model, and our proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo 等NeurIPS 2025 · 被引用 528 次
- SQL-R1: Training Natural Language to SQL Reasoning Model By Reinforcement LearningPeixian Ma, Xialie Zhuang, Chengjin Xu, Xuhui Jiang 等NeurIPS 2025 · 被引用 94 次
- HybridFlow: A Flexible and Efficient RLHF FrameworkGuangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu 等EuroSys 2025 · 被引用 61 次
- MuMA-ToM: Multi-modal Multi-Agent Theory of MindHaojun Shi, Suyu Ye, Xinyu Fang, Chuanyang Jin 等AAAI 2025 · 被引用 48 次
- WorldGPT: Empowering LLM as Multimodal World ModelZhiqi Ge, Hongzhe Huang, Mingze Zhou, Juncheng Li 等ACM MM 2024 · 被引用 35 次
相关 Paper
- MVP: Enhancing Video Large Language Models via Self-supervised Masked Video PredictionXiaokun Sun, Zezhong Wu, Zewen Ding, Linli XuACL 2026 · 被引用 1 次
- EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language ModelsYiyang Fang, Wenke Huang, Pei Fu, Yihao Yang 等CVPR 2026 · 被引用 4 次
- Theory of Mind for Multi-Agent Collaboration via Large Language ModelsHuao Li, Yu Quan Chong, Simon Stepputtis, Joseph Campbell 等EMNLP 2023 · 被引用 57 次
- TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement LearningTao Wu, Li Yang, Gen Zhan, Yabin ZHANG 等CVPR 2026 · 被引用 7 次
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoTBaoqi Pei, Yifei Huang, Jilan Xu, Yuping He 等NeurIPS 2025 · 被引用 21 次
