ScratchNet: A Multi-modal Benchmark for Evaluating and Advancing LLMs on Scratch Programming Tasks
Yuan Si, Simeng Han, Daming Li, Hanyuan Shi, Jialu Zhang
摘要
Large language models (LLMs) have achieved impressive performance on text-based programming tasks, yet they remain unreliable for block-based languages such as Scratch. Scratch programs feature deeply nested, nonlinear structures, event-driven concurrency across multiple sprites, and tight coupling between code and multimedia assets---properties that differ fundamentally from textual code. Consequently, LLMs frequently misinterpret Scratch semantics and propose large, invasive edits that are syntactically valid but semantically misaligned when repairing buggy programs. We introduce ScratchNet, the first executable benchmark designed to systematically evaluate and advance LLM-based repair for Scratch programs. The benchmark comprises 100 carefully curated projects from the public Scratch repository, each selected for high structural and semantic complexity. Every project is paired with an executable test suite, a bug description and corresponding fix, block-level edit constraints that define a minimal semantically correct repair, and the multimedia assets required for faithful execution. We construct the benchmark through a human-in-the-loop pipeline that combines automated project mining with expert validation of trigger--mechanism--outcome semantics and representative bug patterns, with particular emphasis on event ordering, concurrency, and state management. To enable rigorous and reproducible evaluation, we propose a three-layer executable protocol that measures functional correctness through VM-level execution, repair quality through block-level edit distance and behavioral trajectory comparisons, and explanation quality through structured rubrics. Using this benchmark, we study project and bug understanding, trigger and mechanism identification, functional repair, and the effect of lightweight domain adaptation. ScratchNet establishes a reproducible foundation and a closed-loop framework for evaluating and post-training LLMs on block-based programming tasks.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- VisionScratch: LLM-Based Automated Feedback Generation using Code-Produced Videos for Scratch ProgramsYuan Si, Daming Li, Hanyuan Shi, Jialu ZhangFSE 2026 · 被引用 1 次
- On the Applicability of Language Models to Block-Based ProgramsElisabeth Griebl, Benedikt Fein, Florian Obermüller, Gordon Fraser 等ICSE 2023 · 被引用 5 次
- Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification InferenceThanh Le-Cong, Bach Le, Toby MurrayACL 2025
- XCodeEval: An Execution-based Large Scale Multilingual Multitask Benchmark for Code Understanding, Generation, Translation and RetrievalMohammad Abdullah Matin Khan, M. Saiful Bari, Xuan Do Long, Weishi Wang 等ACL 2024 · 被引用 21 次
- LLM-based Agents for Automated Bug Fixing: How Far Are We?Xiangxin Meng, Zexiong Ma, Pengfei Gao, Chao PengICSE 2026
