Lune

ISSTA2026顶会

ScratchNet: A Multi-modal Benchmark for Evaluating and Advancing LLMs on Scratch Programming Tasks

Yuan Si, Simeng Han, Daming Li, Hanyuan Shi, Jialu Zhang

2026年份

摘要

Large language models (LLMs) have achieved impressive performance on text-based programming tasks, yet they remain unreliable for block-based languages such as Scratch. Scratch programs feature deeply nested, nonlinear structures, event-driven concurrency across multiple sprites, and tight coupling between code and multimedia assets---properties that differ fundamentally from textual code. Consequently, LLMs frequently misinterpret Scratch semantics and propose large, invasive edits that are syntactically valid but semantically misaligned when repairing buggy programs. We introduce ScratchNet, the first executable benchmark designed to systematically evaluate and advance LLM-based repair for Scratch programs. The benchmark comprises 100 carefully curated projects from the public Scratch repository, each selected for high structural and semantic complexity. Every project is paired with an executable test suite, a bug description and corresponding fix, block-level edit constraints that define a minimal semantically correct repair, and the multimedia assets required for faithful execution. We construct the benchmark through a human-in-the-loop pipeline that combines automated project mining with expert validation of trigger--mechanism--outcome semantics and representative bug patterns, with particular emphasis on event ordering, concurrency, and state management. To enable rigorous and reproducible evaluation, we propose a three-layer executable protocol that measures functional correctness through VM-level execution, repair quality through block-level edit distance and behavioral trajectory comparisons, and explanation quality through structured rubrics. Using this benchmark, we study project and bug understanding, trigger and mechanism identification, functional repair, and the effect of lightweight domain adaptation. ScratchNet establishes a reproducible foundation and a closed-loop framework for evaluating and post-training LLMs on block-based programming tasks.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖