Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models
Haoxiang Sun, Yingqian Min, Zhipeng Chen, Xin Zhao, Ji-Rong Wen
Abstract
The rapid advancement of large reasoning models has saturated existing math benchmarks, underscoring the urgent need for more challenging evaluation frameworks. To address this, we introduce OlymMATH, a rigorously curated, Olympiad-level math benchmark comprising 350 problems, each with parallel English and Chinese versions. OlymMATH is the first benchmark to unify dual evaluation paradigms within a single suite: (1) natural language evaluation through OlymMATH-EASY and OlymMATH-HARD, comprising 200 computational problems with numerical answers for objective rule-based assessment, and (2) formal verification through OlymMATH-LEAN, offering 150 problems formalized in Lean 4 for rigorous process-level evaluation. All problems are manually sourced from printed publications to minimize data contamination, verified by experts, and span four core domains. Extensive experiments reveal the benchmark's significant challenge, and our analysis also uncovers consistent performance gaps between languages and identifies cases where models employ heuristic "guessing" rather than rigorous reasoning. To further support community research, we release 582k+ reasoning trajectories, a visualization tool, and expert solutions at https://github.com/RUCAIBox/OlymMATH .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- Beyond Pass@ 1: Self-Play with Variational Problem Synthesis Sustains RLVRXiao Liang, Zhong-Zhi Li, Yeyun Gong, Yelong Shen et al.ICLR 2026 · 57 citations
- Optimizing Length Compression in Large Reasoning ModelsZhengxiang Cheng, Dongping Chen, Mingyang Fu, Tianyi ZhouACL 2026 · 32 citations
- Nudging the Boundaries of LLM ReasoningJustin Chih-Yao Chen, Xiangyu Peng, Prafulla Kumar Choubey, Kung-Hsiang Huang et al.ICLR 2026 · 25 citations
- GeoGramBench: Benchmarking the Geometric Program Reasoning in Modern LLMsShixian Luo, Zhu zezhou, Yu Yuan, Yuncheng Yang et al.ICLR 2026 · 15 citations
- Don't Just Fine-tune the Agent, Tune the EnvironmentSiyuan Lu, Zechuan Wang, Hongxuan Zhang, Qintong Wu et al.ICLR 2026 · 13 citations
Builds on6
- Goedel-Prover-V2: Scaling Formal Theorem Proving with Scaffolded Data Synthesis and Self-CorrectionYong Lin, Shange Tang, Bohan Lyu, Ziran Yang et al.ICLR 2026 · 160 citations
- Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language ModelsDaman Arora, Himanshu Gaurav Singh, MausamEMNLP 2023 · 36 citations
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific ProblemsChaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu et al.ACL 2024 · 18 citations
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark for Large Language ModelsBofei Gao, Feifan Song, Zhe Yang, Zefan Cai et al.ICLR 2025 · 3 citations
- Large Language and Reasoning Models are Shallow Disjunctive ReasonersIrtaza Khalid, Amir Masoud Nourollah, Steven SchockaertACL 2025
Related papers
- MathNet: A Global Multimodal Benchmark for Mathematical Reasoning and RetrievalShaden Alshammari, Kevin Wen, Abrar Zainal, Mark Hamilton et al.ICLR 2026 · 9 citations
- RBench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning EvaluationMeng-Hao Guo, Jiajun Xu, Yi Zhang, Jiaxi Song et al.ICML 2025
- Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier MathShrey Pandit, Austin Xu, Xuan-Phi Nguyen, Yifei Ming et al.ACL 2026 · 13 citations
- UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language ModelsXin Xu, Jiaxin Zhang, Tianhao Chen, Zitong Chao et al.ICLR 2025
- ProcessBench: Identifying Process Errors in Mathematical ReasoningChujie Zheng, Zhenru Zhang, Beichen Zhang, Runji Lin et al.ACL 2025 · 209 citations
