HARDMath: A Benchmark Dataset for Challenging Problems in Applied Mathematics
Jingxuan Fan, Sarah Martinson, Erik Y. Wang, Kaylie Hausknecht, Jonah Brenner, Danxian Liu, Nianli Peng, Corey Wang, Michael P. Brenner
Abstract
Advanced applied mathematics problems are underrepresented in existing Large Language Model (LLM) benchmark datasets. To address this, we introduce HARDMath, a dataset inspired by a graduate course on asymptotic methods, featuring challenging applied mathematics problems that require analytical approximation techniques. These problems demand a combination of mathematical reasoning, computational tools, and subjective judgment, making them difficult for LLMs. Our framework auto-generates a large number of problems with solutions validated against numerical ground truths. We evaluate both open- and closed-source LLMs on HARDMath-mini, a sub-sampled test set of 366 problems, as well as on 40 word problems formulated in applied science contexts. Even leading closed-source models like GPT-4 achieve only 43.8% overall accuracy with few-shot Chain-of-Thought prompting, and all models demonstrate significantly lower performance compared to results on existing mathematics benchmark datasets. We additionally conduct a detailed error analysis to gain insights into the failure cases of LLMs. These results demonstrate limitations of current LLM performance on advanced graduate-level applied math problems and underscore the importance of datasets like HARDMath to advance mathematical abilities of LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 71855de5-fefa-45e5-b0e4-bdad156449ecCited by top-tier papers4
- FATE: A Formal Benchmark Series for Frontier Algebra of Multiple Difficulty LevelsJiedong Jiang, Wanyi He, Yuefeng Wang, Guoxiong Gao et al.ICLR 2026 · 26 citations
- CriticLean: Critic-Guided Reinforcement Learning for Mathematical FormalizationZhongyuan Peng, Yifan Yao, Kaijing Ma, Shuyue Guo et al.ACL 2026 · 15 citations
- CMT-Benchmark: A Benchmark for Condensed Matter Theory Built by Expert ResearchersHaining Pan, James V. Roggeveen, Erez Berg, Juan Alvarez et al.ICLR 2026 · 6 citations
- Towards Robust Mathematical ReasoningThang Luong, Dawsen Hwang, Hoang H. Nguyen, Golnaz Ghiasi et al.EMNLP 2025
Builds on3
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language ModelsDaman Arora, Himanshu Gaurav Singh, MausamEMNLP 2023 · 36 citations
Related papers
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark for Large Language ModelsBofei Gao, Feifan Song, Zhe Yang, Zefan Cai et al.ICLR 2025 · 3 citations
- Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical ReasoningJoykirat Singh, Akshay Uttama Nambi, Vibhav VineetACL 2025 · 10 citations
- It Ain't Over: A Multi-aspect Diverse Math Word Problem DatasetJiwoo Kim, Youngbin Kim, Ilwoong Baek, JinYeong Bak et al.EMNLP 2023 · 2 citations
- DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Financial DocumentsYilun Zhao, Yitao Long, Hongjun Liu, Ryo Kamoi et al.ACL 2024 · 8 citations
- UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language ModelsXin Xu, Jiaxin Zhang, Tianhao Chen, Zitong Chao et al.ICLR 2025
