ScratchNet: A Multi-modal Benchmark for Evaluating and Advancing LLMs on Scratch Programming Tasks
Yuan Si, Simeng Han, Daming Li, Hanyuan Shi, Jialu Zhang
Abstract
Large language models (LLMs) have achieved impressive performance on text-based programming tasks, yet they remain unreliable for block-based languages such as Scratch. Scratch programs feature deeply nested, nonlinear structures, event-driven concurrency across multiple sprites, and tight coupling between code and multimedia assets---properties that differ fundamentally from textual code. Consequently, LLMs frequently misinterpret Scratch semantics and propose large, invasive edits that are syntactically valid but semantically misaligned when repairing buggy programs. We introduce ScratchNet, the first executable benchmark designed to systematically evaluate and advance LLM-based repair for Scratch programs. The benchmark comprises 100 carefully curated projects from the public Scratch repository, each selected for high structural and semantic complexity. Every project is paired with an executable test suite, a bug description and corresponding fix, block-level edit constraints that define a minimal semantically correct repair, and the multimedia assets required for faithful execution. We construct the benchmark through a human-in-the-loop pipeline that combines automated project mining with expert validation of trigger--mechanism--outcome semantics and representative bug patterns, with particular emphasis on event ordering, concurrency, and state management. To enable rigorous and reproducible evaluation, we propose a three-layer executable protocol that measures functional correctness through VM-level execution, repair quality through block-level edit distance and behavioral trajectory comparisons, and explanation quality through structured rubrics. Using this benchmark, we study project and bug understanding, trigger and mechanism identification, functional repair, and the effect of lightweight domain adaptation. ScratchNet establishes a reproducible foundation and a closed-loop framework for evaluating and post-training LLMs on block-based programming tasks.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- VisionScratch: LLM-Based Automated Feedback Generation using Code-Produced Videos for Scratch ProgramsYuan Si, Daming Li, Hanyuan Shi, Jialu ZhangFSE 2026 · 1 citation
- On the Applicability of Language Models to Block-Based ProgramsElisabeth Griebl, Benedikt Fein, Florian Obermüller, Gordon Fraser et al.ICSE 2023 · 5 citations
- Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification InferenceThanh Le-Cong, Bach Le, Toby MurrayACL 2025
- XCodeEval: An Execution-based Large Scale Multilingual Multitask Benchmark for Code Understanding, Generation, Translation and RetrievalMohammad Abdullah Matin Khan, M. Saiful Bari, Xuan Do Long, Weishi Wang et al.ACL 2024 · 21 citations
- LLM-based Agents for Automated Bug Fixing: How Far Are We?Xiangxin Meng, Zexiong Ma, Pengfei Gao, Chao PengICSE 2026
