Beyond Single-View Sufficiency: CVBench for Cross-View Human Understanding
Tianchen Guo, Chen Liu, Xin Yu
摘要
Human perception of social environments is inherently a multi-view synthesis problem, requiring the integration of complementary and often occluded information across space and time. However, existing benchmarks for Multimodal Large Language Models (MLLMs) are overwhelmingly predicated on a "sufficient-view" assumption, rewarding single-view pattern recognition while failing to evaluate cross-view fusion. To address this critical gap, we introduce CVBench, a large-scale, multi-task benchmark for cross-view human understanding. CVBench comprises 3,000 challenging questions across 12 spatial and temporal tasks, where every item is designed with verifiable single-view insufficiency, mandating that models synthesize disparate evidence to resolve ambiguities. Our comprehensive evaluation of state-of-the-art open and closed-source MLLMs (from InternVL to Gemini 2.5 Pro) reveals a substantial performance gap, with the best models (e.g., Gemini 2.5 Pro, 42% spatial accuracy) falling nearly 50 points behind human performance (94%). We identify a systemic failure mechanism across all models: a dominant "Single-View Bias," whereby models ignore conflicting evidence and default to the most confident but incorrect single-view prediction. This demonstrates that current MLLMs lack the fundamental mechanisms for geometric grounding, identity persistence, and true spatio-temporal fusion. CVBench provides a rigorous diagnostic framework to catalyze the development of next-generation, cross-view–aware architectures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- DeepMultiCap: Performance Capture of Multiple Characters Using Sparse Multiview CamerasYang Zheng, Ruizhi Shao, Yuxiang Zhang, Tao Yu 等ICCV 2021 · 被引用 112 次
- EgoHumans: An Egocentric 3D Multi-Human BenchmarkRawal Khirodkar, Aayush Bansal, Lingni Ma, Richard A. Newcombe 等ICCV 2023 · 被引用 59 次
- Self-supervised Multi-view Multi-Human Association and TrackingYiyang Gan, Ruize Han, Liqiang Yin, Wei Feng 等ACM MM 2021 · 被引用 44 次
- Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMsChun-Hsiao Yeh, Chenyu Wang, Shengbang Tong, Ta Ying Cheng 等AAAI 2026 · 被引用 35 次
相关 Paper
- CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language ModelsJingyao Li, Jingyun Wang, Molin Tan, Haochen Wang 等AAAI 2026
- XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language ModelsXingrui Wang, Jiang Liu, Chao Huang, Xiaodong Yu 等ICLR 2026 · 被引用 4 次
- HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized BenchmarksTing Zhou, Daoyuan Chen, Qirui Jiao, Bolin Ding 等CVPR 2026
- IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMsDavid Ma, Yuanxing Zhang, Jincheng Ren, Jiawei Guo 等ICLR 2026 · 被引用 5 次
- MuirBench: A Comprehensive Benchmark for Robust Multi-image UnderstandingFei Wang, Xingyu Fu, James Y. Huang, Zekun Li 等ICLR 2025
