EduDiag: A Benchmark for Educational Diagnostic Reasoning with Error Tracing and Correction on Large Multimodal Models
Jiali Chen, Yuqi Xue, Xusen Hei, DingBa Fu, Yuancheng Wei, Jiayuan Xie, Yi Cai
摘要
Large multimodal models (LMMs) have achieved impressive performance on multimodal reasoning, becoming crucial technology for the advancement of intelligent questionanswering systems. In real-world educational scenarios, effective teaching extends far beyond providing answers.
Experienced teachers analyze students' incorrect answers to trace underlying errors and provide corrective feedback, termed educational diagnostic reasoning, a capability that remains under-explored in existing LMMs. To bridge this research gap, we introduce Edudiag benchmark, requiring LMMs to reconstruct erroneous reasoning chains from incorrect answers and generate corrective feedback. Through an AI-assisted annotation pipeline with rigorous human verification, we create 8K erroneous reasoning chains and corresponding feedback, spanning three representative educational domains: commonsense, science, and mathematics. Extensive evaluation across 24 leading LMMs highlights Edudiag as a challenging testbed, where even leading proprietary LMMs struggle on it and supervised finetuning (SFT) on open-source LMMs achieves marginal performance gains. Moreover, we conduct analysis experiments and identify three critical insights for educational diagnostic reasoning: (i) Effective error tracing remains the primary bottleneck, while SFT models still fail to reversely identify errors that commonly occur. (ii) Group relative policy optimization (GRPO) mitigates this bottleneck and boosts performance. (iii) LMMs optimized with GRPO can generate plausible yet challenging distractors for multiplechoice questions based on their self-constructed erroneous reasoning chains. We believe Edudiag provides a new direction for evaluating the advanced LMMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning DistractorJiali Chen, Xusen Hei, Yuqi Xue, Yuancheng Wei 等ACM MM 2024 · 被引用 3 次
- MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGIHuanjin Yao, Jiaxing Huang, Yawen Qiu, Michael K. Chen 等ICCV 2025 · 被引用 4 次
- MC-Search: Evaluating and Enhancing Multimodal Agentic Search with Structured Long Reasoning ChainsXuying Ning, Dongqi Fu, Tianxin Wei, Mengting Ai 等ICLR 2026 · 被引用 14 次
- Benchmarking Fine-Grained Error Detection in Multimodal ReasoningChi-Min Chan, Han Zhu, Chunyang Jiang, Jiaming Ji 等ACL 2026
- Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive ReasoningYian Li, Wentao Tian, Yang Jiao, Tianwen Qian 等ACM MM 2025 · 被引用 16 次
