On the naturalness of hardware descriptions
Jaeseong Lee, Pengyu Nie, Junyi Jessy Li, Milos Gligoric
Abstract
Mining software repositories (MSR) has been shown effective for extracting data used to improve various software engineering tasks, including code completion, code repair, code search, and code summarization. Despite a large body of work on MSR, researchers have focused almost exclusively on repositories that contain code written in imperative programming languages, such as Java and C/C++. Unlike prior work, in this paper, we focus on mining publicly available hardware descriptions (HDs) written in hardware description languages (HDLs), such as VHDL. HDLs have unique syntax and semantics compared to popular imperative languages, and learning-based tools available to hardware designers are well behind those used in other application domains. We assembled large HD corpora consisting of source code written in several HDLs and report on their characteristics. Our language model evaluation reveals that HDs possess a high level of naturalness similar to software written in imperative languages. Further, by utilizing our corpora, we built several deep learning models for automated code completion in VHDL; our models take into account unique characteristics of HDLs, including similarities of nearby concurrent signal assignment statements, in-built concurrency, and the frequently used signal types. These characteristics led to more effective neural models, achieving a BLEU score of 37.3, an 8-14-point improvement over rule-based and neural baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2f8bb045-2a90-49be-9221-6eeba78307e3Cited by top-tier papers3
- Learning Deep Semantics for Test CompletionPengyu Nie, Rahul Banerjee, Junyi Jessy Li, Raymond J. Mooney et al.ICSE 2023 · 45 citations
- CrystalBLEU: Precisely and Efficiently Measuring the Similarity of CodeAryaz Eghbali, Michael PradelASE 2022 · 33 citations
- Pitfalls in Experiments with DNN4SE: An Analysis of the State of the PracticeSira Vegas, Sebastian G. ElbaumFSE 2023 · 4 citations
Builds on3
- TreeGen: A Tree-Based Transformer Architecture for Code GenerationZeyu Sun, Qihao Zhu, Yingfei Xiong, Yican Sun et al.AAAI 2020 · 196 citations
- Associating Natural Language Comment and Source Code EntitiesSheena Panthaplackel, Milos Gligoric, Raymond J. Mooney, Junyi Jessy LiAAAI 2020 · 21 citations
- Learning to Update Natural Language Comments Based on Code ChangesSheena Panthaplackel, Pengyu Nie, Milos Gligoric, Junyi Jessy Li et al.ACL 2020 · 1 citation
Related papers
- DeepRTL: Bridging Verilog Understanding and Generation with a Unified Representation ModelYi Liu, Changran Xu, Yunhao Zhou, Zeju Li et al.ICLR 2025
- CirFix: automatically repairing defects in hardware design codeHammad Ahmad, Yu Huang, Westley WeimerASPLOS 2022 · 23 citations
- VerilogASTBench: Benchmark Construction of Verilog AST Dataset with Dual-Stage AST Semantic Enhancement FrameworkLuping Zhang, Chao Chen, Dapeng Yan, Hui Xu et al.FSE 2026
- CraftRTL: High-quality Synthetic Data Generation for Verilog Code Models with Correct-by-Construction Non-Textual Representations and Targeted Code RepairMingjie Liu, Yun-Da Tsai, Wenfei Zhou, Haoxing RenICLR 2025
- VerilogLAVD: LLM-Aided Pattern Generation for Verilog CWE DetectionXiang Long, Yingjie Xia, Li Kuang, Yao Wan et al.ACL 2026
