Beyond Single-View Sufficiency: CVBench for Cross-View Human Understanding
Tianchen Guo, Chen Liu, Xin Yu
Abstract
Human perception of social environments is inherently a multi-view synthesis problem, requiring the integration of complementary and often occluded information across space and time. However, existing benchmarks for Multimodal Large Language Models (MLLMs) are overwhelmingly predicated on a "sufficient-view" assumption, rewarding single-view pattern recognition while failing to evaluate cross-view fusion. To address this critical gap, we introduce CVBench, a large-scale, multi-task benchmark for cross-view human understanding. CVBench comprises 3,000 challenging questions across 12 spatial and temporal tasks, where every item is designed with verifiable single-view insufficiency, mandating that models synthesize disparate evidence to resolve ambiguities. Our comprehensive evaluation of state-of-the-art open and closed-source MLLMs (from InternVL to Gemini 2.5 Pro) reveals a substantial performance gap, with the best models (e.g., Gemini 2.5 Pro, 42% spatial accuracy) falling nearly 50 points behind human performance (94%). We identify a systemic failure mechanism across all models: a dominant "Single-View Bias," whereby models ignore conflicting evidence and default to the most confident but incorrect single-view prediction. This demonstrates that current MLLMs lack the fundamental mechanisms for geometric grounding, identity persistence, and true spatio-temporal fusion. CVBench provides a rigorous diagnostic framework to catalyze the development of next-generation, cross-view–aware architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b67af44d-ef41-4859-b3f4-cb9daf75ae86Builds on22
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- DeepMultiCap: Performance Capture of Multiple Characters Using Sparse Multiview CamerasYang Zheng, Ruizhi Shao, Yuxiang Zhang, Tao Yu et al.ICCV 2021 · 112 citations
- EgoHumans: An Egocentric 3D Multi-Human BenchmarkRawal Khirodkar, Aayush Bansal, Lingni Ma, Richard A. Newcombe et al.ICCV 2023 · 59 citations
- Self-supervised Multi-view Multi-Human Association and TrackingYiyang Gan, Ruize Han, Liqiang Yin, Wei Feng et al.ACM MM 2021 · 44 citations
- Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMsChun-Hsiao Yeh, Chenyu Wang, Shengbang Tong, Ta Ying Cheng et al.AAAI 2026 · 35 citations
Related papers
- CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language ModelsJingyao Li, Jingyun Wang, Molin Tan, Haochen Wang et al.AAAI 2026
- XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language ModelsXingrui Wang, Jiang Liu, Chao Huang, Xiaodong Yu et al.ICLR 2026 · 4 citations
- HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized BenchmarksTing Zhou, Daoyuan Chen, Qirui Jiao, Bolin Ding et al.CVPR 2026
- IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMsDavid Ma, Yuanxing Zhang, Jincheng Ren, Jiawei Guo et al.ICLR 2026 · 5 citations
- MuirBench: A Comprehensive Benchmark for Robust Multi-image UnderstandingFei Wang, Xingyu Fu, James Y. Huang, Zekun Li et al.ICLR 2025
