SketchMind: A Multi-Agent Cognitive Framework for Assessing Student-Drawn Scientific Sketches
Ehsan Latif, Zirak Khan, Xiaoming Zhai
Abstract
Scientific sketches (e.g., models) offer a powerful lens into students' conceptual understanding, yet AI-powered automated assessment of such free-form, visually diverse artifacts remains a critical challenge. Existing solutions often treat sketch evaluation as either an image classification task or monolithic vision-language models, which lack interpretability, pedagogical alignment, and adaptability across cognitive levels. To address these limitations, we present SKETCHMIND, a cognitively grounded, multi-agent framework for evaluating and improving studentdrawn scientific sketches. SKETCHMIND introduces Sketch Reasoning Graphs (SRGs), semantic graph representations that embed domain concepts and Bloom's taxonomy-based cognitive labels. The system comprises modular agents responsible for rubric parsing, sketch perception, cognitive alignment, and iterative feedback with sketch modification, enabling personalized and transparent evaluation. We evaluate SKETCHMIND on a curated dataset of 3,575 student-generated sketches across six science assessment items with different highest order of Bloom's level that require students to draw models to explain phenomena. Compared to baseline GPT-4o performance without SRG (average accuracy: 55.6%), and with bSRG integration achieves 77.1% average accuracy (+21.4% average absolute gain). We also demonstrate that multi-agent orchestration with SRG enhances SKETCHMIND performance, for example, a SketchMind with GPT-4.1 gains an average 8.9% increase in sketch prediction accuracy, outperforming single-agent pipelines across all items. Human evaluators rated the feedback and co-created sketches generated by SKETCHMIND with GPT-4.1, which achieved an average of 4.1 out of 5, significantly higher than those of baseline models (e.g., 2.3 for GPT-4o). Experts noted the system's potential to meaningfully support conceptual growth through guided revision. Our code and (pending approval) dataset will be released to support reproducibility and future research in AI-driven education.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ae59c3a-c49a-4daa-8356-9f857fbcc30fBuilds on5
- SKED: Sketch-guided Text-based 3D EditingAryan Mikaeili, Or Perel, Mehdi Safaee, Daniel Cohen-Or et al.ICCV 2023 · 83 citations
- Sketch2Saliency: Learning to Detect Salient Objects from Human DrawingsAyan Kumar Bhunia, Subhadeep Koley, Amandeep Kumar, Aneeshan Sain et al.CVPR 2023
- ClearSight: Visual Signal Enhancement for Object Hallucination Mitigation in Multimodal Large Language ModelsHao Yin, Guangzong Si, Zilei WangCVPR 2025
- Data-Free Sketch-Based Image RetrievalAbhra Chaudhuri, Ayan Kumar Bhunia, Yi-Zhe Song, Anjan DuttaCVPR 2023
- Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language ModelsJiacong Xu, Shao-Yuan Lo, Bardia Safaei, Vishal M. Patel et al.CVPR 2025
Related papers
- SketchGPT: A Sketch-based Multimodal Interface for Application-Agnostic LLM InteractionZeyuan Huang, Cangjun Gao, Yaxian Shan, Haoxiang Hu et al.UIST 2025 · 8 citations
- EvoSci: A Bio-Inspired Multi-Agent Framework for the Evolution of Scientific DiscoveryXiaoyu Xiong, Yuqi Ren, Deyi XiongACL 2026 · 1 citation
- CSG: Cognitive Structure Generation for Intelligent EducationHengnian Gu, Zhifu Chen, Yuxin Chen, Jin Zhou et al.ICML 2026
- GIVE: Structured Reasoning of Large Language Models with Knowledge Graph Inspired Veracity ExtrapolationJiashu He, Mingyu Derek Ma, Jinxuan Fan, Dan Roth et al.ICML 2025
- Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language ModelsYushi Hu, Weijia Shi, Xingyu Fu, Dan Roth et al.NeurIPS 2024 · 373 citations
