FIXME: Towards End-to-End Benchmarking of LLM-Aided Design Verification
Gwok-Waa Wan, SamZaak Wong, Shengchu Su, Chenxu Niu, Ning Wang, Xinlai Wan, Qixiang Chen, Mengnv Xing, Jingyi Zhang, Jianmin Ye, Yubo Wang, Rongchang Song
摘要
Despite the transformative potential of Large Language Models (LLMs) in hardware design, a comprehensive evaluation of their capabilities in design verification remains underexplored. Current efforts predominantly focus on RTL generation and basic debugging, overlooking the critical domain of functional verification, which is the primary bottleneck in modern design methodologies due to the rapid escalation of hardware complexity. We present FIXME, the first end-to-end, multi-model, and open-source evaluation framework for assessing LLM performance in hardware functional verification (FV) to address this crucial gap. FIXME introduces a structured threelevel difficulty hierarchy spanning six verification sub-domains and 180 diverse tasks, enabling in-depth analysis across the design lifecycle. Leveraging a collaborative AI-human approach, we construct a high-quality dataset using 100% silicon-proven designs, ensuring comprehensive coverage of real-world challenges. Furthermore, we enhance the functional coverage by 45.57% through expert-guided optimization. By rigorously evaluating state-of-the-art LLMs such as GPT-4, Claude3, and LlaMA3, we identify key areas for improvement and outline promising research directions to unlock the full potential of LLM-driven automation in hardware design verification. The benchmark is available at https://github.com/ChatDesignVerification/FIXME .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- QuArch: A Benchmark for Evaluating LLM Reasoning in Computer ArchitectureShvetank Prakash, Andrew Cheng, Mark Mazumder, Arya Tschand 等ICML 2026 · 被引用 3 次
- ChatHLS: Towards Systematic Design Automation and Optimization for High-Level SynthesisRunkai Li, Jia Xiong, Xiuyuan He, Jieru Zhao 等ACL 2026 · 被引用 3 次
它引用的顶会 Paper5
- BetterV: Controlled Verilog Generation with Discriminative GuidanceZehua Pei, Hui-Ling Zhen, Mingxuan Yuan, Yu Huang 等ICML 2024 · 被引用 155 次
- FlexKBQA: A Flexible LLM-Powered Framework for Few-Shot Knowledge Base Question AnsweringZhenyu Li, Sunqi Fan, Yu Gu, Xiuxing Li 等AAAI 2024 · 被引用 143 次
- Towards Developing High Performance RISC-V Processors Using Agile MethodologyYinan Xu, Zihao Yu, Dan Tang, Guokai Chen 等MICRO 2022 · 被引用 108 次
- RTLFixer: Automatically Fixing RTL Syntax Errors with Large Language ModelYunda Tsai, Mingjie Liu, Haoxing RenDAC 2024 · 被引用 95 次
- Let Me Do It For You: Towards LLM Empowered Recommendation via Tool LearningYuyue Zhao, Jiancan Wu, Xiang Wang, Wei Tang 等SIGIR 2024 · 被引用 42 次
相关 Paper
- UVLLM: An Automated Universal RTL Verification Framework using LLMsYuchen Hu, Junhao Ye, Ke Xu, Jialin Sun 等DAC 2025 · 被引用 4 次
- Free and Fair Hardware: A Pathway to Copyright Infringement-Free Verilog Generation using LLMsSam Bush, Matthew DeLorenzo, Phat Tieu, Jeyavijayan RajendranDAC 2025 · 被引用 5 次
- Data is all you need: Finetuning LLMs for Chip Design via an Automated design-data augmentation frameworkKaiyan Chang, Kun Wang, Nan Yang, Ying Wang 等DAC 2024 · 被引用 59 次
- ChatCPU: An Agile CPU Design and Verification Platform with LLMXi Wang, Gwok-Waa Wan, Sam-Zaak Wong, Layton Zhang 等DAC 2024 · 被引用 31 次
- AnalogVerifier: A Neuro-Symbolic Framework for Analog Circuit VerificationYanfang Liu, Mingjun Wang, Peng XU, Rongliang Fu 等ICML 2026
