FERMAT: An Alternative to Accuracy for Numerical Reasoning
Jasivan Alex Sivakumar, Nafise Sadat Moosavi
摘要
While pre-trained language models achieve impressive performance on various NLP benchmarks, they still struggle with tasks that require numerical reasoning. Recent advances in improving numerical reasoning are mostly achieved using very large language models that contain billions of parameters and are not accessible to everyone. In addition, numerical reasoning is measured using a single score on existing datasets. As a result, we do not have a clear understanding of the strengths and shortcomings of existing models on different numerical reasoning aspects and therefore, potential ways to improve them apart from scaling them up. Inspired by CheckList (Ribeiro et al., 2020), we introduce a multi-view evaluation set for numerical reasoning in English, called FERMAT. Instead of reporting a single score on a whole dataset, FERMAT evaluates models on various key numerical reasoning aspects such as number understanding, mathematical operations, and training dependency. Apart from providing a comprehensive evaluation of models on different numerical reasoning aspects, FERMAT enables a systematic and automated generation of an arbitrarily large training or evaluation set for each aspect.The datasets and codes are publicly available to generate further multi-view data for ulterior tasks and languages. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Enhancing Quantitative Reasoning Skills of Large Language Models through Dimension PerceptionYuncheng Huang, Qianyu He, Jiaqing Liang, Sihang Jiang 等ICDE 2024 · 被引用 3 次
- GeoNum: Bridging Numerical Continuity and Language Semantics via Geometric EmbeddingShengkai Jin, Tianyu Chen, Chonghan Gao, Jun HanAAAI 2026
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
相关 Paper
- ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question AnsweringZhiyu Chen, Shiyang Li, Charese Smiley, Zhiqiang Ma 等EMNLP 2022 · 被引用 57 次
- Injecting Numerical Reasoning Skills into Language ModelsMor Geva, Ankit Gupta, Jonathan BerantACL 2020 · 被引用 12 次
- FinanceReasoning: Benchmarking Financial Numerical Reasoning More Credible, Comprehensive and ChallengingZichen Tang, Haihong E, Ziyan Ma, Haoyang He 等ACL 2025 · 被引用 17 次
- Number Cookbook: Number Understanding of Language Models and How to Improve ItHaotong Yang, Yi Hu, Shijia Kang, Zhouchen Lin 等ICLR 2025
- Can Vision-Language Models Evaluate Handwritten Math?Oikantik Nath, Hanani Bathina, Mohammed Safi Ur Rahman Khan, Mitesh M. KhapraACL 2025 · 被引用 10 次
