QiMeng-CodeV-R1: Reasoning-Enhanced Verilog Generation
Yaoyu Zhu, Di Huang, Han-Qi Lyu, Xiaoyun Zhang, Chongxiao Li, Wenxuan Shi, Yutong Wu, Jianan Mu, Jinghua Wang, Yang Zhao, Pengwei Jin, Shuyao Cheng
摘要
Large language models (LLMs) trained via reinforcement learning with verifiable reward (RLVR) have achieved breakthroughs on tasks with explicit, automatable verification, such as software programming and mathematical problems. Extending RLVR to electronic design automation (EDA), especially automatically generating hardware description languages (HDLs) like Verilog from natural-language (NL) specifications, however, poses three key challenges: the lack of automated and accurate verification environments, the scarcity of high-quality NL-code pairs, and the prohibitive computation cost of RLVR. To this end, we introduce CodeV-R1, an RLVR framework for training Verilog generation LLMs. First, we develop a rule-based testbench generator that performs robust equivalence checking against golden references. Second, we propose a round-trip data synthesis method that pairs open-source Verilog snippets with LLM-generated NL descriptions, verifies code-NL-code consistency via the generated testbench, and filters out inequivalent examples to yield a high-quality dataset. Third, we employ a two-stage "distillthen-RL" training pipeline: distillation for the cold start of reasoning abilities, followed by adaptive DAPO, our novel RLVR algorithm that can reduce training cost by adaptively adjusting sampling rate. The resulting model, CodeV-R1-7B, achieves 68.6 % and 72.9 % pass@1 on VerilogEval v2 and RTLLM v1.1, respectively, surpassing prior state-of-the-art by 12∼20 %, while even exceeding the performance of 671B DeepSeek-R1 on RTLLM. We have released our model, training code, and dataset to facilitate research in EDA and LLM communities. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- LLM4Cov: Execution-Aware Agentic Learning for High-coverage Testbench GenerationHejia Zhang, Zhongming Yu, Chia-Tung Ho, Mark Ren 等ICML 2026 · 被引用 4 次
- Efficient Skill Grounding via Code Refactoring with Small Language ModelsSera Choi, Wonje Choi, Saehun Chun, Daehee Lee 等ICML 2026
- QiMeng-PRepair: Precise Code Repair via Edit-Aware Reward OptimizationChangxin Ke, Rui Zhang, Jiaming Guo, Yuanbo Wen 等ACL 2026
它引用的顶会 Paper6
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base ModelJingcheng Hu, Yinmin Zhang, Qi Han, Daxin Jiang 等NeurIPS 2025 · 被引用 533 次
- BetterV: Controlled Verilog Generation with Discriminative GuidanceZehua Pei, Hui-Ling Zhen, Mingxuan Yuan, Yu Huang 等ICML 2024 · 被引用 155 次
- HybridFlow: A Flexible and Efficient RLHF FrameworkGuangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu 等EuroSys 2025 · 被引用 61 次
- MAGE: A Multi-Agent Engine for Automated RTL Code GenerationYujie Zhao, Hejia Zhang, Hanxian Huang, Zhongming Yu 等DAC 2025 · 被引用 20 次
相关 Paper
- RTLFixer: Automatically Fixing RTL Syntax Errors with Large Language ModelYunda Tsai, Mingjie Liu, Haoxing RenDAC 2024 · 被引用 95 次
- Powering Verifiable Learning via Automated Evolutionary Data SynthesisHe Du, Bowen Li, Aijun Yang, Siyang He 等ACL 2026
- Data is all you need: Finetuning LLMs for Chip Design via an Automated design-data augmentation frameworkKaiyan Chang, Kun Wang, Nan Yang, Ying Wang 等DAC 2024 · 被引用 59 次
- DeepRTL: Bridging Verilog Understanding and Generation with a Unified Representation ModelYi Liu, Changran Xu, Yunhao Zhou, Zeju Li 等ICLR 2025
- QiMeng-CRUX: Narrowing the Gap Between Natural Language and Verilog via Core Refined Understanding eXpressionLei Huang, Rui Zhang, Jiaming Guo, Yang Zhang 等AAAI 2026 · 被引用 1 次
