MathConstruct: Challenging LLM Reasoning with Constructive Proofs
Mislav Balunovic, Jasper Dekoninck, Nikola Jovanovic, Ivo Petrov, Martin T. Vechev
Abstract
While Large Language Models (LLMs) demonstrate impressive performance in mathematics, existing math benchmarks come with significant limitations. Many focus on problems with fixed ground-truth answers, and are often saturated due to problem simplicity or the viability of guessing or memorization. Crucially, they capture only a narrow subset of relevant math problems. To address this research gap, we introduce MATHCONSTRUCT, a new benchmark of 121 challenging problems sourced from various math competitions, which targets constructive proofs, a widely encountered problem type requiring the construction of mathematical objects with specific properties. These proofs are particularly suitable for LLM evaluation, as solution correctness can be easily verified. Our automated verifiers also enable MATHCONSTRUCT to generate problem variations, used to evaluate robustness. State-of-the-art LLMs solve only 60% of MATHCONSTRUCT problems, highlighting its complexity and importance for LLM evaluation. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e97c9af1-0e00-45c6-a3eb-0d64efde8e43Cited by top-tier papers2
- Mathematical Proof as a Litmus Test: Revealing Failure Modes of Advanced Large Reasoning ModelsDadi Guo, Jiayu Liu, Zhiyuan Fan, Zhitao He et al.ACL 2026 · 17 citations
- A Survey of Deep Learning for Geometry Problem SolvingJianzhe Ma, Wenxuan Wang, Qin JinACL 2026 · 5 citations
Builds on7
- miniF2F: a cross-system benchmark for formal Olympiad-level mathematicsKunhao Zheng, Jesse Michael Han, Stanislas PoluICLR 2022 · 342 citations
- ConStat: Performance-Based Contamination Detection in Large Language ModelsJasper Dekoninck, Mark Niklas Müller, Martin T. VechevNeurIPS 2024 · 40 citations
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific ProblemsChaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu et al.ACL 2024 · 18 citations
- NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity ClassesLizhou Fan, Wenyue Hua, Lingyao Li, Haoyang Ling et al.ACL 2024 · 8 citations
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark for Large Language ModelsBofei Gao, Feifan Song, Zhe Yang, Zefan Cai et al.ICLR 2025 · 3 citations
Related papers
- MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex ProofsAndreas Opedal, Haruki Shirakami, Bernhard Schölkopf, Abulhair Saparov et al.ICLR 2025
- Reliable Fine-Grained Evaluation of Natural Language Math ProofsWenjie Ma, Andrei Cojocaru, Neel Kolhe, Haihan Zhang et al.ICLR 2026 · 14 citations
- HARDMath: A Benchmark Dataset for Challenging Problems in Applied MathematicsJingxuan Fan, Sarah Martinson, Erik Y. Wang, Kaylie Hausknecht et al.ICLR 2025
- Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables QuestionsZijin Hong, Hao Wu, Su Dong, Junnan Dong et al.AAAI 2026 · 5 citations
- Can Large Language Models Win the International Mathematical Games?Alessio Cocchieri, Luca Ragazzi, Giuseppe Tagliavini, Lorenzo Tordi et al.EMNLP 2025
