HiPhO: How Far Are (M)LLMs from Humans in the Latest High School Physics Olympiad Benchmark?
Fangchen Yu, Haiyuan Wan, Qianjia Cheng, Yuchen Zhang, Jiacheng Chen, Fujun Han, Yulun Wu, Junchi Yao, Ruilizhen Hu, Ning Ding, Yu Cheng, Tao Chen
Abstract
Recently, the physical capabilities of (M)LLMs have garnered increasing attention. However, existing benchmarks for physics suffer from two major gaps: they neither provide systematic and up-to-date coverage of real-world physics competitions such as physics Olympiads, nor enable direct performance comparison with humans. To bridge these gaps, we present HIPHO, the first benchmark dedicated to high school physics Olympiads with human-aligned evaluation. Specifically, HIPHO highlights three key innovations. (1) Comprehensive Data: It compiles 13 latest Olympiad exams from 2024-2025, spanning both international and regional competitions, and covering mixed modalities that encompass problems spanning text-only to diagram-based. (2) Professional Evaluation: We adopt official marking schemes to perform fine-grained grading at both the answer and step level, fully aligned with human examiners to ensure high-quality and domain-specific evaluation. (3) Comparison with Human Contestants: We assign gold, silver, and bronze medals to models based on official medal thresholds, thereby enabling direct comparison between (M)LLMs and human contestants. Our large-scale evaluation of 30 state-of-the-art (M)LLMs shows that: across 13 exams, open-source MLLMs mostly remain at or below the bronze level; open-source LLMs show promising progress with multiple golds; closedsource reasoning MLLMs can achieve 6 to 12 gold medals; and most models still have a significant gap from full marks. These results highlight the performance gap between open-source models and top students, the strong reasoning abilities of closed-source models, and the remaining room for improvement. HIPHO, a human-aligned Olympiad benchmark for multimodal physical reasoning, is opensource at https://github.com/SciYu/HiPhO with a public leaderboard at https://phyarena.github.io/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a18e9a4a-4625-4aaf-9f56-2a3fab0fe61bCited by top-tier papers1
Ask how each one uses itBuilds on2
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific ProblemsChaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu et al.ACL 2024 · 18 citations
- UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language ModelsXin Xu, Qiyun Xu, Tong Xiao, Tianhao Chen et al.ICML 2025
Related papers
- MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language ModelsJiacheng Ruan, Dan Jiang, Xian Gao, Ting Liu et al.AAAI 2026 · 3 citations
- Sim2Reason: Solving Physics Olympiad via Reinforcement Learning on Physics SimulatorsMihir Prabhudesai, Aryan Satpathy, Yangmin Li, Zheyang Qin et al.ICML 2026 · 2 citations
- PRISM-Physics: Causal DAG-Based Process Evaluation for Physics ReasoningWanjia Zhao, Qinwei Ma, Jingzhe Shi, Shirley Wu et al.ICLR 2026 · 4 citations
- PhysReason: A Comprehensive Benchmark towards Physics-Based ReasoningXinyu Zhang, Yuxuan Dong, Yanrui Wu, Jiaxing Huang et al.ACL 2025 · 51 citations
- DeepPhy: Benchmarking Agentic VLMs on Physical ReasoningXinrun Xu, Pi Bu, Ye Wang, Börje F. Karlsson et al.AAAI 2026 · 6 citations
