PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning
Xinyu Zhang, Yuxuan Dong, Yanrui Wu, Jiaxing Huang, Chengyou Jia, Basura Fernando, Mike Zheng Shou, Lingling Zhang, Jun Liu
Abstract
Large language models demonstrate remarkable capabilities across various domains, especially mathematics and logic reasoning. However, current evaluations overlook physicsbased reasoning, a complex task requiring physics theorems and constraints. We present PhysReason, a 1,200-problem benchmark comprising knowledge-based (25%) and reasoningbased (75%) problems, where the latter are divided into three difficulty levels (easy, medium, hard). Notably, problems require an average of 8.1 solution steps, with hard problems requiring 15.6, reflecting the complexity of physicsbased reasoning. We propose the Physics Solution Auto Scoring Framework, incorporating efficient answer-level and comprehensive steplevel evaluations. Top-performing models like Deepseek-R1, Gemini-2.0-Flash-Thinking, and o3-mini-high achieve less than 60% on answerlevel evaluation, with performance dropping from knowledge questions (75.11%) to hard problems (31.95%). Through step-level evaluation, we identify four key bottlenecks: Physics Theorem Application, Physics Process Understanding, Calculation, and Physics Condition Analysis. These findings position Phys-Reason as a novel and comprehensive benchmark for evaluating physics-based reasoning capabilities in large language models. Our code and data will be published at https: //dxzxy12138.github.io/PhysReason/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd1bb88d-367e-4328-8629-5418c4627c39Cited by top-tier papers7
- LEXam: Benchmarking Legal Reasoning on 340 Law ExamsYu Fan, Jingwei Ni, Jakob Merane, Yang Tian et al.ICLR 2026 · 56 citations
- CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter PhysicsWeida Wang, Dongchen Huang, Jiatong Li, Tengchao Yang et al.ICLR 2026 · 11 citations
- CoFFT: Chain of Foresight-Focus Thought for Visual Language ModelsXinyu Zhang, Yuxuan Dong, Lingling Zhang, Chengyou Jia et al.NeurIPS 2025 · 7 citations
- Lean4Physics: Comprehensive Reasoning Framework for College-level Physics in Lean4Yuxin Li, Minghao Liu, Ruida Wang, Wenzhao Ji et al.ICLR 2026 · 7 citations
- PRISM-Physics: Causal DAG-Based Process Evaluation for Physics ReasoningWanjia Zhao, Qinwei Ma, Jingzhe Shi, Shirley Wu et al.ICLR 2026 · 4 citations
Builds on12
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh et al.NeurIPS 2024 · 1,019 citations
- SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language ModelsXiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu et al.ICML 2024 · 220 citations
Related papers
- UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language ModelsXin Xu, Qiyun Xu, Tong Xiao, Tianhao Chen et al.ICML 2025
- MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGIHuanjin Yao, Jiaxing Huang, Yawen Qiu, Michael K. Chen et al.ICCV 2025 · 4 citations
- Scientific logicality enriched methodology for LLM reasoning: A practice in physicsZhaoxin Yu, Nan Xu, Kun Chen, Jiahao Zhao et al.ICML 2026
- FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and ChallengingZichen Tang, Haihong E, Ziyan Ma, Haoyang He et al.ACL 2025 · 17 citations
- OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image ReasoningMingxin Huang, Yongxin Shi, Dezhi Peng, Songxuan Lai et al.ICLR 2026 · 28 citations
