Fuzzing the Boundary between Models and Code in Hybrid AI-Enabled Systems
Xinyu Gao, Yang Feng, Yuchen Lu, Zhenqian Liu, Zhenyu Chen, Baowen Xu
摘要
Deep learning (DL) techniques are increasingly integrated into traditional software systems, giving rise to hybrid AI-enabled systems that combine neural models with program logic. While these systems exhibit remarkable capabilities, their complex and heterogeneous architectures pose significant challenges for reliability and testing, particularly in safety-critical domains such as autonomous driving. Existing testing approaches either target traditional code or isolate neural networks, overlooking failures arising from their interactions. In this paper, we present Neude, a lightweight and extensible coverage-guided fuzzing framework specifically designed for hybrid AI-enabled systems. Unlike existing tools, Neude combines observations of program execution and neural model coverage to guide input mutations toward unexplored state spaces, enabling systematic testing of the entire hybrid system. Moreover, Neude employs domain-aware mutation operators coupled with metamorphic relations, allowing automated bug detection without manual assertions. We evaluate Neude on Pylot, a complex autonomous driving system with tightly coupled neural and program components. Experimental results show that Neude uncovers diverse errors, and further analysis reveals how model uncertainty propagates through deterministic program logic to trigger downstream module failures. Our findings highlight the fragility of current hybrid architectures, calling for a paradigm shift from model-centric testing to system-centric quality assurance that accounts for the intricate interplay between neural and procedural components.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- NeuRI: Diversifying DNN Generation via Inductive Rule InferenceJiawei Liu, Jinjun Peng, Yuyao Wang, Lingming ZhangFSE 2023 · 被引用 24 次
- Audee: Automated Testing for Deep Learning FrameworksQianyu Guo, Xiaofei Xie, Yi Li, Xiaoyu Zhang 等ASE 2020 · 被引用 83 次
- RobOT: Robustness-Oriented Testing for Deep Learning SystemsJingyi Wang, Jialuo Chen, Youcheng Sun, Xingjun Ma 等ICSE 2021 · 被引用 62 次
- Large Language Models Are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language ModelsYinlin Deng, Chunqiu Steven Xia, Haoran Peng, Chenyuan Yang 等ISSTA 2023 · 被引用 253 次
- DriveFuzz: Discovering Autonomous Driving Bugs through Driving Quality-Guided FuzzingSeulbae Kim, Major Liu, Junghwan John Rhee, Yuseok Jeon 等CCS 2022 · 被引用 65 次
