A Differential Testing Framework to Identify Critical AV Failures Leveraging Arbitrary Inputs
Trey Woodlief, Carl Hildebrandt, Sebastian G. Elbaum
Abstract
The proliferation of autonomous vehicles (AVs) has made their failures increasingly evident. Testing efforts aimed at identifying the inputs leading to those failures are challenged by the input's long-tail distribution, whose area under the curve is dominated by rare scenarios. We hypothesize that leveraging emerging open-access datasets can accelerate the exploration of long-tail inputs. Having access to diverse inputs, however, is not sufficient to expose failures; an effective test also requires an oracle to distinguish between correct and incorrect behaviors. Current datasets lack such oracles and developing them is notoriously difficult. In response, we propose DIFFTEST4AV, a differential testing framework designed to address the unique challenges of testing AV systems: 1) for any given input, many outputs may be considered acceptable, 2) the long-tail contains an insurmountable number of inputs to explore, and 3) the AV's continuous execution loop requires for failures to persist in order to affect the system. DIFFTEST4AV integrates statistical analysis to identify meaningful behavioral variations, judges their importance in terms of the severity of these differences, and incorporates sequential analysis to detect persistent errors indicative of potential system-level failures. Our study on 5 versions of the commercially-available, road-deployed comma.ai OpenPilot system, using 3 available image datasets, demonstrates the capabilities of the framework to detect high-severity, highconfidence, long-running test failures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f93f7208-adad-4d46-bf8d-77459b722586Builds on4
- Fuzz testing based data augmentation to improve robustness of deep neural networksXiang Gao, Ripon K. Saha, Mukul R. Prasad, Abhik RoychoudhuryICSE 2020 · 116 citations
- Semantic Image Fuzzing of AI Perception SystemsTrey Woodlief, Sebastian G. Elbaum, Kevin SullivanICSE 2022 · 17 citations
- Generating Realistic and Diverse Tests for LiDAR-Based Perception SystemsGarrett Christian, Trey Woodlief, Sebastian G. ElbaumICSE 2023 · 11 citations
- REDriver: Runtime Enforcement for Autonomous VehiclesYang Sun, Christopher M. Poskitt, Xiaodong Zhang, Jun SunICSE 2024 · 5 citations
Related papers
- S3C: Spatial Semantic Scene Coverage for Autonomous VehiclesTrey Woodlief, Felipe Toledo, Sebastian G. Elbaum, Matthew B. DwyerICSE 2024 · 11 citations
- CorrectAD: A Self-Correcting Agentic System to Improve End-to-end Planning in Autonomous DrivingEnhui Ma, Lijun Zhou, Tao Tang, Jiahuan Zhang et al.AAAI 2026
- MOSAT: finding safety violations of autonomous driving systems using multi-objective genetic algorithmHaoxiang Tian, Yan Jiang, Guoquan Wu, Jiren Yan et al.FSE 2022 · 74 citations
- MoDitector: Module-Directed Testing for Autonomous Driving SystemsRenzhi Wang, Mingfei Cheng, Xiaofei Xie, Yuan Zhou et al.ISSTA 2025 · 3 citations
- AIDE: An Automatic Data Engine for Object Detection in Autonomous DrivingMingfu Liang, Jong-Chyi Su, Samuel Schulter, Sparsh Garg et al.CVPR 2024
