Free and Fair Hardware: A Pathway to Copyright Infringement-Free Verilog Generation using LLMs
Sam Bush, Matthew DeLorenzo, Phat Tieu, Jeyavijayan Rajendran
Abstract
Limitations in Large Language Model (LLM) capabilities for hardware design tasks, such as generating functional Verilog codes, have motivated various fine-tuning optimizations utilizing curated hardware datasets from open-source repositories. However, these datasets remain limited in size and contain minimal checks on licensing for reuse, resulting in potential copyright violations by fine-tuned LLMs. Therefore, we propose an evaluation benchmark to estimate the risk of Verilog-trained LLMs to generate copyright-protected codes. To minimize this risk, we present an open-source Verilog dataset, FreeSet, containing over 220k files, along with the automated dataset curation framework utilized to provide additional guarantees of fair-use Verilog data. We then execute an LLM fine-tuning framework consisting of continual pre-training, resulting in a fine-tuned Llama model for Verilog, FreeV. Our results indicate that FreeV demonstrates the smallest risk of copyright-infringement among prior works, with only a 3% violation rate. Furthermore, experimental results demonstrate improvements in Verilog generation functionality over its baseline model, improving VerilogEval pass@10 rates by over 10%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9bbdc19-c5c9-4fce-bcd2-88529f6d3b33Builds on6
- BetterV: Controlled Verilog Generation with Discriminative GuidanceZehua Pei, Hui-Ling Zhen, Mingxuan Yuan, Yu Huang et al.ICML 2024 · 155 citations
- RTLFixer: Automatically Fixing RTL Syntax Errors with Large Language ModelYunda Tsai, Mingjie Liu, Haoxing RenDAC 2024 · 95 citations
- GNN4IP: Graph Neural Network for Hardware Intellectual Property Piracy DetectionRozhin Yasaei, Shih-Yuan Yu, Emad Kasaeyan Naeini, Mohammad Abdullah Al FaruqueDAC 2021 · 43 citations
- Copyright Traps for Large Language ModelsMatthieu Meeus, Igor Shilov, Manuel Faysse, Yves-Alexandre de MontjoyeICML 2024 · 39 citations
- SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text GenerationXiaoze Liu, Ting Sun, Tianyang Xu, Feijie Wu et al.EMNLP 2024 · 4 citations
Related papers
- PyraNet: A Multi-Layered Hierarchical Dataset for VerilogBardia Nadimi, Ghali Omar Boutaib, Hao ZhengDAC 2025 · 8 citations
- CraftRTL: High-quality Synthetic Data Generation for Verilog Code Models with Correct-by-Construction Non-Textual Representations and Targeted Code RepairMingjie Liu, Yun-Da Tsai, Wenfei Zhou, Haoxing RenICLR 2025
- Data is all you need: Finetuning LLMs for Chip Design via an Automated design-data augmentation frameworkKaiyan Chang, Kun Wang, Nan Yang, Ying Wang et al.DAC 2024 · 59 citations
- VerilogASTBench: Benchmark Construction of Verilog AST Dataset with Dual-Stage AST Semantic Enhancement FrameworkLuping Zhang, Chao Chen, Dapeng Yan, Hui Xu et al.FSE 2026
- FIXME: Towards End-to-End Benchmarking of LLM-Aided Design VerificationGwok-Waa Wan, SamZaak Wong, Shengchu Su, Chenxu Niu et al.AAAI 2026
