On the Applicability of Language Models to Block-Based Programs
Elisabeth Griebl, Benedikt Fein, Florian Obermüller, Gordon Fraser, René Just
摘要
Block-based programming languages like Scratch are increasingly popular for programming education and end-user programming. Recent program analyses build on the insight that source code can be modelled using techniques from natural language processing. Many of the regularities of source code that support this approach are due to the syntactic overhead imposed by textual programming languages. This syntactic overhead, however, is precisely what block-based languages remove in order to simplify programming. Consequently, it is unclear how well this modelling approach performs on block-based programming languages. In this paper, we investigate the applicability of language models for the popular block-based programming language Scratch. We model Scratch programs using n-gram models, the most essential type of language model, and transformers, a popular deep learning model. Evaluation on the example tasks of code completion and bug finding confirm that blocks inhibit predictability, but the use of language models is nevertheless feasible. Our findings serve as foundation for improving tooling and analyses for block-based languages.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- Code Prediction by Feeding Trees to TransformersSeohyun Kim, Jinman Zhao, Yuchi Tian, Satish ChandraICSE 2021 · 被引用 179 次
- Multi-task Learning based Pre-trained Language Model for Code CompletionFang Liu, Ge Li, Yunfei Zhao, Zhi JinASE 2020 · 被引用 162 次
- Big code != big vocabulary: open-vocabulary models for source codeRafael-Michael Karampatsis, Hlib Babii, Romain Robbes, Charles Sutton 等ICSE 2020 · 被引用 140 次
- Empirical study of transformers for source codeNadezhda Chirkova, Sergey TroshinFSE 2021 · 被引用 53 次
相关 Paper
- VisionScratch: LLM-Based Automated Feedback Generation using Code-Produced Videos for Scratch ProgramsYuan Si, Daming Li, Hanyuan Shi, Jialu ZhangFSE 2026 · 被引用 1 次
- ScratchNet: A Multi-modal Benchmark for Evaluating and Advancing LLMs on Scratch Programming TasksYuan Si, Simeng Han, Daming Li, Hanyuan Shi 等ISSTA 2026
- Verified from Scratch: Program Analysis for Learners' ProgramsAndreas Stahlbauer, Christoph Frädrich, Gordon FraserASE 2020 · 被引用 7 次
- Emergent Representations of Program Semantics in Language Models Trained on ProgramsCharles Jin, Martin C. RinardICML 2024 · 被引用 34 次
- Can Machines Read Coding Manuals Yet? - A Benchmark for Building Better Language Models for Code UnderstandingIbrahim Abdelaziz, Julian Dolby, Jamie P. McCusker, Kavitha SrinivasAAAI 2022 · 被引用 7 次
