MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
Xuehai He, Weixi Feng, Kaizhi Zheng, Yujie Lu, Wanrong Zhu, Jiachen Li, Yue Fan, Jianfeng Wang, Linjie Li, Zhengyuan Yang, Kevin Lin, William Yang Wang
Abstract
Multimodal Language Language Models (MLLMs) demonstrate the emerging abilities of "world models" -- interpreting and reasoning about complex real-world dynamics. To assess these abilities, we posit videos are the ideal medium, as they encapsulate rich representations of real-world dynamics and causalities. To this end, we introduce MMWorld, a new benchmark for multi-discipline, multi-faceted multimodal video understanding. MMWorld distinguishes itself from previous video understanding benchmarks with two unique advantages: (1) multi-discipline, covering various disciplines that often require domain expertise for comprehensive understanding; (2) multi-faceted reasoning, including explanation, counterfactual thinking, future prediction, etc. MMWorld consists of a human-annotated dataset to evaluate MLLMs with questions about the whole videos and a synthetic dataset to analyze MLLMs within a single modality of perception. Together, MMWorld encompasses 1,910 videos across seven broad disciplines and 69 subdisciplines, complete with 6,627 question-answer pairs and associated captions. The evaluation includes 2 proprietary and 10 open-source MLLMs, which struggle on MMWorld (e.g., GPT-4V performs the best with only 52.3% accuracy), showing large room for improvement. Further ablation studies reveal other interesting findings such as models' different skill sets from humans. We hope MMWorld can serve as an essential step towards world model evaluation in videos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 10af1e8e-dc64-4d21-b0b3-328890f752faCited by top-tier papers10
- VLM4D: Towards Spatiotemporal Awareness in Vision Language ModelsShijie Zhou, Alexander Vilesov, Xuehai He, Ziyu Wan et al.ICCV 2025 · 9 citations
- AMusE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker UnderstandingSanjoy Chowdhury, Karren Dai Yang, Xudong Liu, Fartash Faghri et al.CVPR 2026 · 5 citations
- Multi-Agent System for Comprehensive Soccer UnderstandingJiayuan Rao, Zifeng Li, Haoning Wu, Ya Zhang et al.ACM MM 2025 · 5 citations
- VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative VideosJiashuo Yu, Yue Wu, Meng Chu, Zhifei Ren et al.ICCV 2025 · 3 citations
- MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeXLiuyue Xie, Avik Kuthiala, George Z. Wei, Ce Zheng et al.AAAI 2026 · 1 citation
Builds on20
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMsDavid Ma, Yuanxing Zhang, Jincheng Ren, Jiawei Guo et al.ICLR 2026 · 5 citations
- CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language ModelsJingyao Li, Jingyun Wang, Molin Tan, Haochen Wang et al.AAAI 2026
- MMVU: Measuring Expert-Level Multi-Discipline Video UnderstandingYilun Zhao, Haowei Zhang, Lujing Xie, Tongyan Hu et al.CVPR 2025
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video AnalysisChaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li et al.CVPR 2025
- Beyond Single View: A Comprehensive Benchmark for Medical Multimodal Large Language Models on Multi-Image UnderstandingDexuan Xu, Jiayin Yuan, Jianing Wang, Yanyuan Chen et al.ACL 2026
