CraftRTL: High-quality Synthetic Data Generation for Verilog Code Models with Correct-by-Construction Non-Textual Representations and Targeted Code Repair
Mingjie Liu, Yun-Da Tsai, Wenfei Zhou, Haoxing Ren
Abstract
Despite the significant progress made in code generation with large language models, challenges persist, especially with hardware description languages such as Verilog. This paper first presents an analysis of fine-tuned LLMs on Verilog coding, with synthetic data from prior methods. We identify two main issues: difficulties in handling non-textual representations (Karnaugh maps, state-transition diagrams and waveforms) and significant variability during training with models randomly making "minor" mistakes. To address these limitations, we enhance data curation by creating correct-by-construction data targeting non-textual representations. Additionally, we introduce an automated framework that generates error reports from various model checkpoints and injects these errors into opensource code to create targeted code repair data. Our fine-tuned Starcoder2-15B outperforms prior state-of-the-art results by 3.8%, 10.9%, 6.6% for pass@1 on VerilogEval-Machine, VerilogEval-Human, and RTLLM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e06bf977-c9f6-4db8-ac0f-4b3c688adfb6Cited by top-tier papers7
- QiMeng-CodeV-R1: Reasoning-Enhanced Verilog GenerationYaoyu Zhu, Di Huang, Han-Qi Lyu, Xiaoyun Zhang et al.NeurIPS 2025 · 46 citations
- ChipSeek: Optimizing Verilog Generation via EDA-Integrated Reinforcement LearningZhirong Chen, Kaiyan Chang, Zhuolin Li, Cangyuan Li et al.ACL 2026 · 6 citations
- Free and Fair Hardware: A Pathway to Copyright Infringement-Free Verilog Generation using LLMsSam Bush, Matthew DeLorenzo, Phat Tieu, Jeyavijayan RajendranDAC 2025 · 5 citations
- OSIRIS: Bridging Analog Circuit Design and Machine Learning with Scalable Dataset GenerationGiuseppe Chiari, Michele Piccoli, Davide ZoniICLR 2026 · 4 citations
- QiMeng-CRUX: Narrowing the Gap Between Natural Language and Verilog via Core Refined Understanding eXpressionLei Huang, Rui Zhang, Jiaming Guo, Yang Zhang et al.AAAI 2026 · 1 citation
Builds on18
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Model Alignment as Prospect Theoretic OptimizationKawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky et al.ICML 2024 · 973 citations
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun et al.ICLR 2024 · 945 citations
Related papers
- PyraNet: A Multi-Layered Hierarchical Dataset for VerilogBardia Nadimi, Ghali Omar Boutaib, Hao ZhengDAC 2025 · 8 citations
- VerilogASTBench: Benchmark Construction of Verilog AST Dataset with Dual-Stage AST Semantic Enhancement FrameworkLuping Zhang, Chao Chen, Dapeng Yan, Hui Xu et al.FSE 2026
- Data is all you need: Finetuning LLMs for Chip Design via an Automated design-data augmentation frameworkKaiyan Chang, Kun Wang, Nan Yang, Ying Wang et al.DAC 2024 · 59 citations
- BetterV: Controlled Verilog Generation with Discriminative GuidanceZehua Pei, Hui-Ling Zhen, Mingxuan Yuan, Yu Huang et al.ICML 2024 · 155 citations
- RTLFixer: Automatically Fixing RTL Syntax Errors with Large Language ModelYunda Tsai, Mingjie Liu, Haoxing RenDAC 2024 · 95 citations
