FIXME: Towards End-to-End Benchmarking of LLM-Aided Design Verification
Gwok-Waa Wan, SamZaak Wong, Shengchu Su, Chenxu Niu, Ning Wang, Xinlai Wan, Qixiang Chen, Mengnv Xing, Jingyi Zhang, Jianmin Ye, Yubo Wang, Rongchang Song
Abstract
Despite the transformative potential of Large Language Models (LLMs) in hardware design, a comprehensive evaluation of their capabilities in design verification remains underexplored. Current efforts predominantly focus on RTL generation and basic debugging, overlooking the critical domain of functional verification, which is the primary bottleneck in modern design methodologies due to the rapid escalation of hardware complexity. We present FIXME, the first end-to-end, multi-model, and open-source evaluation framework for assessing LLM performance in hardware functional verification (FV) to address this crucial gap. FIXME introduces a structured threelevel difficulty hierarchy spanning six verification sub-domains and 180 diverse tasks, enabling in-depth analysis across the design lifecycle. Leveraging a collaborative AI-human approach, we construct a high-quality dataset using 100% silicon-proven designs, ensuring comprehensive coverage of real-world challenges. Furthermore, we enhance the functional coverage by 45.57% through expert-guided optimization. By rigorously evaluating state-of-the-art LLMs such as GPT-4, Claude3, and LlaMA3, we identify key areas for improvement and outline promising research directions to unlock the full potential of LLM-driven automation in hardware design verification. The benchmark is available at https://github.com/ChatDesignVerification/FIXME .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- QuArch: A Benchmark for Evaluating LLM Reasoning in Computer ArchitectureShvetank Prakash, Andrew Cheng, Mark Mazumder, Arya Tschand et al.ICML 2026 · 3 citations
- ChatHLS: Towards Systematic Design Automation and Optimization for High-Level SynthesisRunkai Li, Jia Xiong, Xiuyuan He, Jieru Zhao et al.ACL 2026 · 3 citations
Builds on5
- BetterV: Controlled Verilog Generation with Discriminative GuidanceZehua Pei, Hui-Ling Zhen, Mingxuan Yuan, Yu Huang et al.ICML 2024 · 155 citations
- FlexKBQA: A Flexible LLM-Powered Framework for Few-Shot Knowledge Base Question AnsweringZhenyu Li, Sunqi Fan, Yu Gu, Xiuxing Li et al.AAAI 2024 · 143 citations
- Towards Developing High Performance RISC-V Processors Using Agile MethodologyYinan Xu, Zihao Yu, Dan Tang, Guokai Chen et al.MICRO 2022 · 108 citations
- RTLFixer: Automatically Fixing RTL Syntax Errors with Large Language ModelYunda Tsai, Mingjie Liu, Haoxing RenDAC 2024 · 95 citations
- Let Me Do It For You: Towards LLM Empowered Recommendation via Tool LearningYuyue Zhao, Jiancan Wu, Xiang Wang, Wei Tang et al.SIGIR 2024 · 42 citations
Related papers
- UVLLM: An Automated Universal RTL Verification Framework using LLMsYuchen Hu, Junhao Ye, Ke Xu, Jialin Sun et al.DAC 2025 · 4 citations
- Free and Fair Hardware: A Pathway to Copyright Infringement-Free Verilog Generation using LLMsSam Bush, Matthew DeLorenzo, Phat Tieu, Jeyavijayan RajendranDAC 2025 · 5 citations
- Data is all you need: Finetuning LLMs for Chip Design via an Automated design-data augmentation frameworkKaiyan Chang, Kun Wang, Nan Yang, Ying Wang et al.DAC 2024 · 59 citations
- ChatCPU: An Agile CPU Design and Verification Platform with LLMXi Wang, Gwok-Waa Wan, Sam-Zaak Wong, Layton Zhang et al.DAC 2024 · 31 citations
- AnalogVerifier: A Neuro-Symbolic Framework for Analog Circuit VerificationYanfang Liu, Mingjun Wang, Peng XU, Rongliang Fu et al.ICML 2026
