Fuzzing the Boundary between Models and Code in Hybrid AI-Enabled Systems
Xinyu Gao, Yang Feng, Yuchen Lu, Zhenqian Liu, Zhenyu Chen, Baowen Xu
Abstract
Deep learning (DL) techniques are increasingly integrated into traditional software systems, giving rise to hybrid AI-enabled systems that combine neural models with program logic. While these systems exhibit remarkable capabilities, their complex and heterogeneous architectures pose significant challenges for reliability and testing, particularly in safety-critical domains such as autonomous driving. Existing testing approaches either target traditional code or isolate neural networks, overlooking failures arising from their interactions. In this paper, we present Neude, a lightweight and extensible coverage-guided fuzzing framework specifically designed for hybrid AI-enabled systems. Unlike existing tools, Neude combines observations of program execution and neural model coverage to guide input mutations toward unexplored state spaces, enabling systematic testing of the entire hybrid system. Moreover, Neude employs domain-aware mutation operators coupled with metamorphic relations, allowing automated bug detection without manual assertions. We evaluate Neude on Pylot, a complex autonomous driving system with tightly coupled neural and program components. Experimental results show that Neude uncovers diverse errors, and further analysis reveals how model uncertainty propagates through deterministic program logic to trigger downstream module failures. Our findings highlight the fragility of current hybrid architectures, calling for a paradigm shift from model-centric testing to system-centric quality assurance that accounts for the intricate interplay between neural and procedural components.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get fbe2c963-55fc-4138-9739-481625f6f870Related papers
- NeuRI: Diversifying DNN Generation via Inductive Rule InferenceJiawei Liu, Jinjun Peng, Yuyao Wang, Lingming ZhangFSE 2023 · 24 citations
- Audee: Automated Testing for Deep Learning FrameworksQianyu Guo, Xiaofei Xie, Yi Li, Xiaoyu Zhang et al.ASE 2020 · 83 citations
- RobOT: Robustness-Oriented Testing for Deep Learning SystemsJingyi Wang, Jialuo Chen, Youcheng Sun, Xingjun Ma et al.ICSE 2021 · 62 citations
- Large Language Models Are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language ModelsYinlin Deng, Chunqiu Steven Xia, Haoran Peng, Chenyuan Yang et al.ISSTA 2023 · 253 citations
- DriveFuzz: Discovering Autonomous Driving Bugs through Driving Quality-Guided FuzzingSeulbae Kim, Major Liu, Junghwan John Rhee, Yuseok Jeon et al.CCS 2022 · 65 citations
