FERMAT: An Alternative to Accuracy for Numerical Reasoning
Jasivan Alex Sivakumar, Nafise Sadat Moosavi
Abstract
While pre-trained language models achieve impressive performance on various NLP benchmarks, they still struggle with tasks that require numerical reasoning. Recent advances in improving numerical reasoning are mostly achieved using very large language models that contain billions of parameters and are not accessible to everyone. In addition, numerical reasoning is measured using a single score on existing datasets. As a result, we do not have a clear understanding of the strengths and shortcomings of existing models on different numerical reasoning aspects and therefore, potential ways to improve them apart from scaling them up. Inspired by CheckList (Ribeiro et al., 2020), we introduce a multi-view evaluation set for numerical reasoning in English, called FERMAT. Instead of reporting a single score on a whole dataset, FERMAT evaluates models on various key numerical reasoning aspects such as number understanding, mathematical operations, and training dependency. Apart from providing a comprehensive evaluation of models on different numerical reasoning aspects, FERMAT enables a systematic and automated generation of an arbitrarily large training or evaluation set for each aspect.The datasets and codes are publicly available to generate further multi-view data for ulterior tasks and languages. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 26c754a0-8d3a-417f-99b8-c7a22c8701eeCited by top-tier papers2
- Enhancing Quantitative Reasoning Skills of Large Language Models through Dimension PerceptionYuncheng Huang, Qianyu He, Jiaqing Liang, Sihang Jiang et al.ICDE 2024 · 3 citations
- GeoNum: Bridging Numerical Continuity and Language Semantics via Geometric EmbeddingShengkai Jin, Tianyu Chen, Chonghan Gao, Jun HanAAAI 2026
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
Related papers
- ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question AnsweringZhiyu Chen, Shiyang Li, Charese Smiley, Zhiqiang Ma et al.EMNLP 2022 · 57 citations
- Injecting Numerical Reasoning Skills into Language ModelsMor Geva, Ankit Gupta, Jonathan BerantACL 2020 · 12 citations
- FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and ChallengingZichen Tang, Haihong E, Ziyan Ma, Haoyang He et al.ACL 2025 · 17 citations
- Number Cookbook: Number Understanding of Language Models and How to Improve ItHaotong Yang, Yi Hu, Shijia Kang, Zhouchen Lin et al.ICLR 2025
- Can Vision-Language Models Evaluate Handwritten Math?Oikantik Nath, Hanani Bathina, Mohammed Safi Ur Rahman Khan, Mitesh M. KhapraACL 2025 · 10 citations
