On the Applicability of Language Models to Block-Based Programs
Elisabeth Griebl, Benedikt Fein, Florian Obermüller, Gordon Fraser, René Just
Abstract
Block-based programming languages like Scratch are increasingly popular for programming education and end-user programming. Recent program analyses build on the insight that source code can be modelled using techniques from natural language processing. Many of the regularities of source code that support this approach are due to the syntactic overhead imposed by textual programming languages. This syntactic overhead, however, is precisely what block-based languages remove in order to simplify programming. Consequently, it is unclear how well this modelling approach performs on block-based programming languages. In this paper, we investigate the applicability of language models for the popular block-based programming language Scratch. We model Scratch programs using n-gram models, the most essential type of language model, and transformers, a popular deep learning model. Evaluation on the example tasks of code completion and bug finding confirm that blocks inhibit predictability, but the use of language models is nevertheless feasible. Our findings serve as foundation for improving tooling and analyses for block-based languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a69289f7-cce3-4362-9cd7-774a0fc67dacCited by top-tier papers1
Ask how each one uses itBuilds on5
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng et al.ICLR 2021 · 1,644 citations
- Code Prediction by Feeding Trees to TransformersSeohyun Kim, Jinman Zhao, Yuchi Tian, Satish ChandraICSE 2021 · 179 citations
- Multi-task Learning based Pre-trained Language Model for Code CompletionFang Liu, Ge Li, Yunfei Zhao, Zhi JinASE 2020 · 162 citations
- Big code != big vocabulary: open-vocabulary models for source codeRafael-Michael Karampatsis, Hlib Babii, Romain Robbes, Charles Sutton et al.ICSE 2020 · 140 citations
- Empirical study of transformers for source codeNadezhda Chirkova, Sergey TroshinFSE 2021 · 53 citations
Related papers
- VisionScratch: LLM-Based Automated Feedback Generation using Code-Produced Videos for Scratch ProgramsYuan Si, Daming Li, Hanyuan Shi, Jialu ZhangFSE 2026 · 1 citation
- ScratchNet: A Multi-modal Benchmark for Evaluating and Advancing LLMs on Scratch Programming TasksYuan Si, Simeng Han, Daming Li, Hanyuan Shi et al.ISSTA 2026
- Verified from Scratch: Program Analysis for Learners' ProgramsAndreas Stahlbauer, Christoph Frädrich, Gordon FraserASE 2020 · 7 citations
- Emergent Representations of Program Semantics in Language Models Trained on ProgramsCharles Jin, Martin C. RinardICML 2024 · 34 citations
- Can Machines Read Coding Manuals Yet? - A Benchmark for Building Better Language Models for Code UnderstandingIbrahim Abdelaziz, Julian Dolby, Jamie P. McCusker, Kavitha SrinivasAAAI 2022 · 7 citations
