HiPhO: How Far Are (M)LLMs from Humans in the Latest High School Physics Olympiad Benchmark?
Fangchen Yu, Haiyuan Wan, Qianjia Cheng, Yuchen Zhang, Jiacheng Chen, Fujun Han, Yulun Wu, Junchi Yao, Ruilizhen Hu, Ning Ding, Yu Cheng, Tao Chen
摘要
Recently, the physical capabilities of (M)LLMs have garnered increasing attention. However, existing benchmarks for physics suffer from two major gaps: they neither provide systematic and up-to-date coverage of real-world physics competitions such as physics Olympiads, nor enable direct performance comparison with humans. To bridge these gaps, we present HIPHO, the first benchmark dedicated to high school physics Olympiads with human-aligned evaluation. Specifically, HIPHO highlights three key innovations. (1) Comprehensive Data: It compiles 13 latest Olympiad exams from 2024-2025, spanning both international and regional competitions, and covering mixed modalities that encompass problems spanning text-only to diagram-based. (2) Professional Evaluation: We adopt official marking schemes to perform fine-grained grading at both the answer and step level, fully aligned with human examiners to ensure high-quality and domain-specific evaluation. (3) Comparison with Human Contestants: We assign gold, silver, and bronze medals to models based on official medal thresholds, thereby enabling direct comparison between (M)LLMs and human contestants. Our large-scale evaluation of 30 state-of-the-art (M)LLMs shows that: across 13 exams, open-source MLLMs mostly remain at or below the bronze level; open-source LLMs show promising progress with multiple golds; closedsource reasoning MLLMs can achieve 6 to 12 gold medals; and most models still have a significant gap from full marks. These results highlight the performance gap between open-source models and top students, the strong reasoning abilities of closed-source models, and the remaining room for improvement. HIPHO, a human-aligned Olympiad benchmark for multimodal physical reasoning, is opensource at https://github.com/SciYu/HiPhO with a public leaderboard at https://phyarena.github.io/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper2
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific ProblemsChaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu 等ACL 2024 · 被引用 18 次
- UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language ModelsXin Xu, Qiyun Xu, Tong Xiao, Tianhao Chen 等ICML 2025
相关 Paper
- MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language ModelsJiacheng Ruan, Dan Jiang, Xian Gao, Ting Liu 等AAAI 2026 · 被引用 3 次
- Sim2Reason: Solving Physics Olympiad via Reinforcement Learning on Physics SimulatorsMihir Prabhudesai, Aryan Satpathy, Yangmin Li, Zheyang Qin 等ICML 2026 · 被引用 2 次
- PRISM-Physics: Causal DAG-Based Process Evaluation for Physics ReasoningWanjia Zhao, Qinwei Ma, Jingzhe Shi, Shirley Wu 等ICLR 2026 · 被引用 4 次
- PhysReason: A Comprehensive Benchmark towards Physics-Based ReasoningXinyu Zhang, Yuxuan Dong, Yanrui Wu, Jiaxing Huang 等ACL 2025 · 被引用 51 次
- DeepPhy: Benchmarking Agentic VLMs on Physical ReasoningXinrun Xu, Pi Bu, Ye Wang, Börje F. Karlsson 等AAAI 2026 · 被引用 6 次
