Towards Robust Mathematical Reasoning
Thang Luong, Dawsen Hwang, Hoang H. Nguyen, Golnaz Ghiasi, Yuri Chervonyi, Insuk Seo, Junsu Kim, Garrett Bingham, Jonathan Lee, Swaroop Mishra, Alex Zhai, Clara Huiyi Hu
摘要
Finding the right north-star metrics is highly critical for advancing mathematical reasoning capabilities of foundation models, especially given that existing evaluations are either too easy or only focusing on getting correct short answers. To address these issues, we present IMO-Bench, a suite of advanced reasoning benchmarks that specifically targets the level of the International Mathematical Olympiad (IMO), the most prestigious venue for young mathematicians. IMO-AnswerBench first tests models on 400 diverse Olympiad problems with verifiable short answers. IMO-ProofBench is the next-level evaluation for proof-writing capabilities, which includes both basic and advanced IMO problems as well as detailed grading guidelines to facilitate automatic grading. These benchmarks played a crucial role in our historic achievement of the gold-level performance at IMO 2025 with Gemini Deep Think (Luong and Lockhart, 2025). Our model achieved 80.0% on IMO-AnswerBench and 65.7% on the advanced IMO-ProofBench, surpassing the best non-Gemini models by large margins of 6.9% and 42.4% respectively. We also showed that autograders built with Gemini reasoning correlate well with human evaluations and construct IMO-GradingBench, with 1000 human gradings on proofs, to enable further progress in automatic evaluation of long-form answers. We hope that IMO-Bench will help the community towards advancing robust mathematical reasoning and release it at https://imobench . github.io. * If Correct: Correct * If Incorrect: Incorrect CRITICAL CONSTRAINT: Do not add any text, explanations, or formatting outside the <thinking> tags or the final output. Output exmaple: <thinking> 1. Golden Answer: (-∞, -4) ∪ (-4, ∞) 2. Extracted Model Answer: ∅ (the empty set) * Evaluation Categories: The expected output must be one of the following categories: 'correct', 'partial', 'almost', 'incorrect', or 'not found'. * Score Identification: The extraction is based on identifying the keyword used by the evaluator to summarize their conclusion. The criteria associated with these keywords are: * incorrect: The evaluator concluded that the solution is completely incorrect or irrelevant. * partial: The evaluator concluded that the solution is partially correct but has significant errors or omissions. * almost: The evaluator concluded that the solution is almost correct but contains minor errors or inaccuracies. * correct: The evaluator concluded that the solution is fully correct and complete. * not_found: The evaluation response does not clearly contain one of the four explicit scores listed above. * Extraction: Determine the provided score from the response and extract the category ( 'correct', 'partial', 'almost', or 'incorrect'). If a score cannot be reliably identified within the text, the output must be 'not_found'.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Does Your Reasoning Model Implicitly Know When to Stop Thinking?Zixuan Huang, Xin Xia, Yuxi Ren, Jianbin Zheng 等ICML 2026 · 被引用 21 次
- PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated ReasoningJingcheng Hu, Yinmin Zhang, Shijie Shang, Xiaobo Yang 等ACL 2026 · 被引用 15 次
- InT: Self-Proposed Interventions Enable Credit Assignment in LLM ReasoningMatthew Y. R. Yang, Hao Bai, Ian Wu, Gene Yang 等ICLR 2026 · 被引用 12 次
- MathNet: A Global Multimodal Benchmark for Mathematical Reasoning and RetrievalShaden Alshammari, Kevin Wen, Abrar Zainal, Mark Hamilton 等ICLR 2026 · 被引用 9 次
- Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RLIan Wu, Yuxiao Qu, Amrith Setlur, Aviral KumarICML 2026 · 被引用 8 次
它引用的顶会 Paper3
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- miniF2F: a cross-system benchmark for formal Olympiad-level mathematicsKunhao Zheng, Jesse Michael Han, Stanislas PoluICLR 2022 · 被引用 342 次
- HARDMath: A Benchmark Dataset for Challenging Problems in Applied MathematicsJingxuan Fan, Sarah Martinson, Erik Y. Wang, Kaylie Hausknecht 等ICLR 2025
相关 Paper
- Reliable Fine-Grained Evaluation of Natural Language Math ProofsWenjie Ma, Andrei Cojocaru, Neel Kolhe, Haihan Zhang 等ICLR 2026 · 被引用 14 次
- QEDBench: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical ProofsSantiago Gonzalez, Alireza Amiribavandpour, Peter Ye, Edward Zhang 等ICML 2026 · 被引用 1 次
- Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier MathShrey Pandit, Austin Xu, Xuan-Phi Nguyen, Yifei Ming 等ACL 2026 · 被引用 13 次
- OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language ModelsQiguang Chen, Chengyu Luan, Jiajun Wu, Qiming Yu 等ACL 2026 · 被引用 1 次
- RefineBench: Evaluating Refinement Capability of Language Models via ChecklistsYoung-Jun Lee, Seungone Kim, Byung-Kwan Lee, Minkyeong Moon 等ICLR 2026 · 被引用 13 次
