LWDIFF: an LLM-Assisted Differential Testing Framework for Webassembly Runtimes
Shiyao Zhou, Jincheng Wang, He Ye, Hao Zhou, Claire Le Goues, Xiapu Luo
Abstract
WebAssembly (Wasm) runtimes execute Wasm programs, a popular low-level language for efficiently executing high-level languages in browsers, with broad applications across diverse domains. The correctness of those runtimes is critical for both functionality and security of Wasm execution, motivating testing approaches that target Wasm runtimes specifically. However, existing Wasm testing frameworks fail to generate test cases that effectively test all three phases of runtime, i.e., decoding, validation, and execution. To address this research gap, we propose a new differential testing framework for Wasm runtimes, which leverages knowledge from the Wasm language specification that prior techniques overlooked, enhancing comprehensive testing of runtime functionality. Specifically, we first use a large language model to extract that knowledge from the specification. We use that knowledge in the context of multiple novel mutation operators that generate test cases with diverse features to test all three runtime phases. We evaluate LWDIFF by applying it to eight Wasm runtimes. Compared with the state-of-the-art Wasm testers, LWDIFF achieves the highest branch coverage and identifies the largest number of bugs. In total, LWDIFF discovers 31 bugs across eight runtimes, all of which are confirmed, with 25 of them previously undiscovered.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd89b5b4-c498-4efd-b584-c2dbcfc617d2Cited by top-tier papers2
- When Specifications Meet Reality: Uncovering API Inconsistencies in Ethereum InfrastructureJie Ma, Ningyu He, Jinwen Xi, Mingzhe Xing et al.OOPSLA 2026
- Debugging Performance Issues in WebAssembly Runtimes via Mutation-based InferenceRuiying Zeng, Shuyao Jiang, Wenxuan Zhao, Yangfan ZhouICSE 2026
Builds on17
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Coverage-based Greybox Fuzzing as Markov ChainMarcel Böhme, Van-Thuan Pham, Abhik RoychoudhuryCCS 2016 · 1,026 citations
- One Fuzzing Strategy to Rule Them AllMingyuan Wu, Ling Jiang, Jiahong Xiang, Yanwei Huang et al.ICSE 2022 · 64 citations
- LLMIF: Augmented Large Language Model for Fuzzing IoT DevicesJincheng Wang, Le Yu, Xiapu LuoS&P 2024 · 61 citations
- A fast in-place interpreter for WebAssemblyBen L. TitzerOOPSLA 2022 · 44 citations
Related papers
- WADIFF: A Differential Testing Framework for WebAssembly RuntimesShiyao Zhou, Muhui Jiang, Weimin Chen, Hao Zhou et al.ASE 2023 · 14 citations
- WASIT: Deep and Continuous Differential Testing of WebAssembly System Interface ImplementationsYage Hu, Wen Zhang, Botang Xiao, Qingchen Kong et al.SOSP 2025
- WASMaker: Differential Testing of WebAssembly Runtimes via Semantic-Aware Binary GenerationShangtong Cao, Ningyu He, Xinyu She, Yixuan Zhang et al.ISSTA 2024 · 8 citations
- WEST: Specification-Based Test Generation for WebAssemblyDongjun Youn, Wonho Shin, Sukyoung RyuASE 2025
- WASCII: Bridging WebAssembly Specifications and Implementations through LLM-Enhanced ValidationYeqi Fu, Kaihang Ji, Yuanpeng Wang, Zong Cao et al.ISSTA 2026
