Can Vision-Language Models Evaluate Handwritten Math?
Oikantik Nath, Hanani Bathina, Mohammed Safi Ur Rahman Khan, Mitesh M. Khapra
Abstract
Recent advancements in Vision-Language Models (VLMs) have opened new possibilities in automatic grading of handwritten student responses, particularly in mathematics. However, a comprehensive study to test the ability of VLMs to evaluate and reason over handwritten content remains absent. To address this gap, we introduce FERMAT, a benchmark designed to assess the ability of VLMs to detect, localize and correct errors in handwritten mathematical content. FERMAT spans four key error dimensions - computational, conceptual, notational, and presentation - and comprises over 2,200 handwritten math solutions derived from 609 manually curated problems from grades 7-12 with intentionally introduced perturbations. Using FERMAT we benchmark nine VLMs across three tasks: error detection, localization, and correction. Our results reveal significant shortcomings in current VLMs in reasoning over handwritten text, with Gemini-1.5-Pro achieving the highest error correction rate (77%). We also observed that some models struggle with processing handwritten content, as their accuracy improves when handwritten inputs are replaced with printed text or images. These findings highlight the limitations of current VLMs and reveal new avenues for improvement. We release FERMAT and all the associated resources in the open-source to drive further research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 89417b89-89ad-4364-8d9c-b0f11e9bdf10Cited by top-tier papers4
- Uni-MuMER: Unified Multi-Task Fine-Tuning of Vision-Language Model for Handwritten Mathematical Expression RecognitionYu Li, Jin Jiang, Jianhua Zhu, Shuai Peng et al.NeurIPS 2025 · 7 citations
- AmIWrite: Exploring Scalable One-on-One Handwriting-Based Tutoring for Mathematical Problem-Solving with an LLM-Powered AI TutorZiyi Liu, Yuzhao Chen, Haoyu Ji, Runlin Duan et al.CHI 2026 · 1 citation
- VEHME: A Vision-Language Model For Evaluating Handwritten Mathematics ExpressionsThu Phuong Nguyen, Duc M. Nguyen, Hyotaek Jeon, Hyunwook Lee et al.EMNLP 2025 · 1 citation
- Seeing Symbols, Missing Structure: A Real-World Handwritten Mathematical Expression Recognition Benchmark for Large ModelsSheng Jiang, Lin Zhu, Runrui Li, Mei Wang et al.ICML 2026
Builds on2
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific ProblemsChaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu et al.ACL 2024 · 18 citations
Related papers
- MMErroR: A Benchmark for Erroneous Reasoning in Vision-Language ModelsYang Shi, Yifeng Xie, Minzhe Guo, Liangsi Lu et al.ACL 2026 · 9 citations
- DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language ModelsChengke Zou, Xingang Guo, Rui Yang, Junyu Zhang et al.ICLR 2025
- FERMAT: An Alternative to Accuracy for Numerical ReasoningJasivan Alex Sivakumar, Nafise Sadat MoosaviACL 2023 · 2 citations
- OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language ModelsQiguang Chen, Chengyu Luan, Jiajun Wu, Qiming Yu et al.ACL 2026 · 1 citation
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
