IRFuzzer: Specialized Fuzzing for LLVM Backend Code Generation
Yuyang Rong, Zhanghan Yu, Zhenkai Weng, Stephen Neuendorffer, Hao Chen
Abstract
Modern compilers, such as LLVM, are complex. Due to their complexity, manual testing is unlikely to suffice, yet formal verification is difficult to scale. End-to-end fuzzing can be used, but it has difficulties in discovering LLVM backend problems for two reasons. First, frontend preprocessing and middle optimization shield the backend from seeing diverse inputs. Second, branch coverage cannot provide effective feedback as LLVM backend contains much reusable code. In this paper, we implement IRFuzzer to investigate the need of specialized fuzzing of the LLVM compiler backend. We focus on two approaches to improve the fuzzer: guaranteed input validity using constrained mutations to improve input diversity and new metrics to improve feedback quality. The mutator in IRFuzzer can generate a wide range of LLVM IR inputs, including structured control flow, vector types, and function definitions. The system instruments coding patterns in the compiler to monitor the execution status of instruction selection. The instrumentation not only provides new coverage feedback on the matcher table but also guides the mutator on architecture-specific intrinsics. We ran IRFuzzer on 29 mature LLVM backend targets. IRFuzzer discovered 78 new, confirmed bugs in LLVM upstream, none of which existing fuzzers could discover. This demonstrates that IRFuzzer is far more effective than existing fuzzers. Upon receiving our bug report, the developers have fixed 57 bugs and back-ported five fixes to LLVM 15, which shows that specialized fuzzing provides actionable insights to LLVM developers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 62466069-5895-48ec-83cc-4d6a9b2bd100Cited by top-tier papers3
- Translation Validation for LLVM's AArch64 BackendRyan Berger, Mitch Briles, Nader Boushehrinejad Moradi, Nicholas Coughlin et al.OOPSLA 2025 · 3 citations
- CLIR: Liveness-Driven and Structure-Aware Fuzzing for the Cranelift CompilerShangtong Cao, Tianlei Song, Qiuping Yi, Tianyu Chen et al.ISSTA 2026
- Semantic Reification: A New Paradigm for Random Program GenerationKavya Chopra, Cong Li, Thodoris Sotiropoulos, Zhendong SuPLDI 2026
Builds on31
- Coverage-based Greybox Fuzzing as Markov ChainMarcel Böhme, Van-Thuan Pham, Abhik RoychoudhuryCCS 2016 · 1,026 citations
- Evaluating Fuzz TestingGeorge Klees, Andrew Ruef, Benji Cooper, Shiyi Wei et al.CCS 2018 · 753 citations
- Angora: Efficient Fuzzing by Principled SearchPeng Chen, Hao ChenS&P 2018 · 616 citations
- CollAFL: Path Sensitive FuzzingShuitao Gan, Chao Zhang, Xiaojun Qin, Xuwen Tu et al.S&P 2018 · 426 citations
- REDQUEEN: Fuzzing with Input-to-State CorrespondenceCornelius Aschermann, Sergej Schumilo, Tim Blazytko, Robert Gawlik et al.NDSS 2019 · 413 citations
Related papers
- FLUX: Finding Bugs with LLVM IR Based Unit Test CrossoversEric Liu, Shengjie Xu, David LieASE 2023 · 8 citations
- GrayC: Greybox Fuzzing of Compilers and Analysers for CKarine Even-Mendoza, Arindam Sharma, Alastair F. Donaldson, Cristian CadarISSTA 2023 · 52 citations
- Optimization-Directed Compiler Fuzzing for Continuous Translation ValidationJaeseong Kwon, Bongjun Jang, Juneyoung Lee, Kihong HeoPLDI 2025 · 5 citations
- One Engine to Fuzz 'em All: Generic Language Processor Testing with Semantic ValidationYongheng Chen, Rui Zhong, Hong Hu, Hangfan Zhang et al.S&P 2021 · 70 citations
- Alive2: bounded translation validation for LLVMNuno P. Lopes, Juneyoung Lee, Chung-Kil Hur, Zhengyang Liu et al.PLDI 2021 · 109 citations
