Towards Reliable AI: Adequacy Metrics for Ensuring the Quality of System-level Testing of Autonomous Vehicles
Neelofar, Aldeida Aleti
摘要
AI-powered systems have gained widespread popularity in various domains, including Autonomous Vehicles (AVs). However, ensuring their reliability and safety is challenging due to their complex nature. Conventional test adequacy metrics, designed to evaluate the effectiveness of traditional software testing, are often insufficient or impractical for these systems. White-box metrics, which are specifically designed for these systems, leverage neuron coverage information. These coverage metrics necessitate access to the underlying AI model and training data, which may not always be available. Furthermore, the existing adequacy metrics exhibit weak correlations with the ability to detect faults in the generated test suite, creating a gap that we aim to bridge in this study. In this paper, we introduce a set of black-box test adequacy metrics called "Test suite Instance Space Adequacy" (TISA) metrics, which can be used to gauge the effectiveness of a test suite. The TISA metrics offer a way to assess both the diversity and coverage of the test suite and the range of bugs detected during testing. Additionally, we introduce a framework that permits testers to visualise the diversity and coverage of the test suite in a two-dimensional space, facilitating the identification of areas that require improvement. We evaluate the efficacy of the TISA metrics by examining their correlation with the number of bugs detected in system-level simulation testing of AVs. A strong correlation, coupled with the short computation time, indicates their effectiveness and efficiency in estimating the adequacy of testing AVs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- DeepGini: prioritizing massive tests to enhance the robustness of deep neural networksYang Feng, Qingkai Shi, Xinyu Gao, Jun Wan 等ISSTA 2020 · 被引用 206 次
- Model-based exploration of the frontier of behaviours for deep learning system testingVincenzo Riccio, Paolo TonellaFSE 2020 · 被引用 134 次
- DeepHyperion: exploring the feature space of deep learning-based systems through illumination searchTahereh Zohdinasab, Vincenzo Riccio, Alessio Gambi, Paolo TonellaISSTA 2021 · 被引用 76 次
- MOSAT: finding safety violations of autonomous driving systems using multi-objective genetic algorithmHaoxiang Tian, Yan Jiang, Guoquan Wu, Jiren Yan 等FSE 2022 · 被引用 74 次
- Importance-driven deep learning system testingSimos Gerasimou, Hasan Ferit Eniser, Alper Sen, Alper ÇakanICSE 2020 · 被引用 65 次
相关 Paper
- PhysCov: Physical Test Coverage for Autonomous VehiclesCarl Hildebrandt, Meriel von Stein, Sebastian G. ElbaumISSTA 2023 · 被引用 15 次
- Measuring and Mitigating Gaps in Structural TestingSoneya Binta Hossain, Matthew B. Dwyer, Sebastian G. Elbaum, Anh Nguyen-TuongICSE 2023 · 被引用 7 次
- RobOT: Robustness-Oriented Testing for Deep Learning SystemsJingyi Wang, Jialuo Chen, Youcheng Sun, Xingjun Ma 等ICSE 2021 · 被引用 62 次
- Fuzzing the Boundary between Models and Code in Hybrid AI-Enabled SystemsXinyu Gao, Yang Feng, Yuchen Lu, Zhenqian Liu 等ISSTA 2026
- How Does Simulation-Based Testing for Self-Driving Cars Match Human Perception?Christian Birchler, Tanzil Kombarabettu Mohammed, Pooja Rani, Teodora Nechita 等FSE 2024 · 被引用 21 次
