Towards Reliable AI: Adequacy Metrics for Ensuring the Quality of System-level Testing of Autonomous Vehicles
Neelofar, Aldeida Aleti
Abstract
AI-powered systems have gained widespread popularity in various domains, including Autonomous Vehicles (AVs). However, ensuring their reliability and safety is challenging due to their complex nature. Conventional test adequacy metrics, designed to evaluate the effectiveness of traditional software testing, are often insufficient or impractical for these systems. White-box metrics, which are specifically designed for these systems, leverage neuron coverage information. These coverage metrics necessitate access to the underlying AI model and training data, which may not always be available. Furthermore, the existing adequacy metrics exhibit weak correlations with the ability to detect faults in the generated test suite, creating a gap that we aim to bridge in this study. In this paper, we introduce a set of black-box test adequacy metrics called "Test suite Instance Space Adequacy" (TISA) metrics, which can be used to gauge the effectiveness of a test suite. The TISA metrics offer a way to assess both the diversity and coverage of the test suite and the range of bugs detected during testing. Additionally, we introduce a framework that permits testers to visualise the diversity and coverage of the test suite in a two-dimensional space, facilitating the identification of areas that require improvement. We evaluate the efficacy of the TISA metrics by examining their correlation with the number of bugs detected in system-level simulation testing of AVs. A strong correlation, coupled with the short computation time, indicates their effectiveness and efficiency in estimating the adequacy of testing AVs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f7931f36-fc42-434a-b259-fcbe14293cccCited by top-tier papers1
Ask how each one uses itBuilds on6
- DeepGini: prioritizing massive tests to enhance the robustness of deep neural networksYang Feng, Qingkai Shi, Xinyu Gao, Jun Wan et al.ISSTA 2020 · 206 citations
- Model-based exploration of the frontier of behaviours for deep learning system testingVincenzo Riccio, Paolo TonellaFSE 2020 · 134 citations
- DeepHyperion: exploring the feature space of deep learning-based systems through illumination searchTahereh Zohdinasab, Vincenzo Riccio, Alessio Gambi, Paolo TonellaISSTA 2021 · 76 citations
- MOSAT: finding safety violations of autonomous driving systems using multi-objective genetic algorithmHaoxiang Tian, Yan Jiang, Guoquan Wu, Jiren Yan et al.FSE 2022 · 74 citations
- Importance-driven deep learning system testingSimos Gerasimou, Hasan Ferit Eniser, Alper Sen, Alper ÇakanICSE 2020 · 65 citations
Related papers
- PhysCov: Physical Test Coverage for Autonomous VehiclesCarl Hildebrandt, Meriel von Stein, Sebastian G. ElbaumISSTA 2023 · 15 citations
- Measuring and Mitigating Gaps in Structural TestingSoneya Binta Hossain, Matthew B. Dwyer, Sebastian G. Elbaum, Anh Nguyen-TuongICSE 2023 · 7 citations
- RobOT: Robustness-Oriented Testing for Deep Learning SystemsJingyi Wang, Jialuo Chen, Youcheng Sun, Xingjun Ma et al.ICSE 2021 · 62 citations
- Fuzzing the Boundary between Models and Code in Hybrid AI-Enabled SystemsXinyu Gao, Yang Feng, Yuchen Lu, Zhenqian Liu et al.ISSTA 2026
- How Does Simulation-Based Testing for Self-Driving Cars Match Human Perception?Christian Birchler, Tanzil Kombarabettu Mohammed, Pooja Rani, Teodora Nechita et al.FSE 2024 · 21 citations
