HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
Keliang Li, Hongze Shen, Hao Shi, Ruibing Hou, Hong Chang, Jie Huang, Chenghao Jia, Wen Wang, Yiling Wu, Dongmei Jiang, Shiguang Shan, Xilin Chen
摘要
The aspiration for artifical general intelligence, fueled by the rapid progress of multimodal models, demands human-comparable performance across diverse environments. We propose HumanPCR, an evaluation suite for probing MLLMs' capacity about human-related visual contexts across three hierarchical levels: Perception, Comprehension, and Reasoning (denoted by Human-P, Human-C, and Human-R, respectively). Human-P and Human-C feature over 6,000 human-verified multiplechoice questions, assessing massive tasks of 9 dimensions, including but not limited to essential skills frequently overlooked by existing benchmarks. Human-R offers a challenging manually curated video reasoning test that requires integrating multiple visual evidences, proactively extracting context beyond question cues, and applying human-like expertise. Each question includes human-annotated Chain-of-Thought (CoT) rationales with key visual evidence to support further research. Extensive evaluations on over 30 state-of-the-art models exhibit significant challenges in human-centric visual understanding, particularly in tasks involving detailed space perception, temporal understanding, and mind modeling. Moreover, analysis of Human-R reveals the struggle of models in extracting essential proactive visual evidence from diverse human scenes and their faulty reliance on query-guided retrieval. Even with advanced techniques like scaling visual contexts and test-time thinking yield only limited benefits. We hope HumanPCR and our findings will advance the development, evaluation, and human-centric application of multimodal models. * Equal contribution. Author order was determined randomly. † Corresponding author. H um an -P Spati ality Spatial Relation Objec t Existe nce Hum an Pres enc e Po st ur e Bo dy Po stu re Ha nd St at e Ha nd -O bj ec t In te ra ct io n Bo dy O rie nt at io n Ga ze Es tim at io n Ap pea ran ce Clo th ing At tri bu te Acc ess ory Rec ogn itio n Bod ypar t Visib ility Physica l Attribu te Contact Human-Object Contact Human-Hum an Contact Huma n Self-C ontac t Ide nti ty Fac e Rec ogn itio n Ide nti ty Clu ste rin g Human-C B eh av io r Ge st ur e Em ot io n Ba si c Ac tio n Kn ow le dg e-Ba se d Ac tio n Proc edu re Se qu en tia l Ac tio n Goa l Pla nni ng Proc edur e Depe nden ce Multiple Human Sequenc ial Action Irrelevant Action Rel atio n Human Comparis on Socia l Rela tion Gro up Act ivit y Str ug gle De tec tio n S ce n e Cr ow d Ev en t Cu ltu ra l Ev
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper37
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu 等NeurIPS 2023 · 被引用 698 次
相关 Paper
- Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language ModelsYuansen Liu, Haiming Tang, Jinlong Peng, Jiangning Zhang 等ICLR 2026 · 被引用 5 次
- X-PCR: A Benchmark for Cross-modality Progressive Clinical Reasoning in Ophthalmic DiagnosisGui Wang, Zehao Zhong, YongSong Zhou, Yudong Li 等CVPR 2026
- Decoding Scientific Experimental Images: The SPUR Benchmark for Perception, Understanding, and ReasoningJunpeng Ding, Zichen Tang, Haihong E, Mengyuan Ji 等ACL 2026
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsWeiye Xu, Jiahao Wang, Weiyun Wang, Zhe Chen 等ICLR 2026 · 被引用 103 次
- MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in VideosKejian Zhu, Zhuoran Jin, Hongbang Yuan, Jiachun Li 等ICLR 2026 · 被引用 22 次
