SAFE: Harnessing LLM for Scenario-Driven ADS Testing from Multimodal Crash Data
Siwei Luo, Yang Zhang, Yao Deng, Linfeng Liang, Xi Zheng
Abstract
Ensuring the safety of Autonomous Driving Systems (ADS) requires realistic and reproducible test scenarios, yet extracting such scenarios from multimodal crash reports remains a major challenge. Large Language Models (LLMs) often hallucinate and lose map structure, resulting in unrealistic road layouts and vehicle behaviors. To address this, we introduce SAFE, a novel Scenario-based ADS testing Framework via multimodal Extraction, which leverages Retrieval-Augmented Generation (RAG), knowledge-grounded prompting, Chain-of-Thought (CoT) reasoning, and self-validation to improve scenario reconstruction from multimodal crash data.
SAFE achieves 93.8% accuracy in extracting road network details, 80.0% for actor information, and 100% for environmental context. In human studies, SAFE outperforms LCTGen and AC3R in reconstructing consistent road networks and vehicle behaviors. Under identical ADS and simulator settings, SAFE detects 39 and 71 more safety violations than LCTGen and AC3R, respectively, and reproduces 12 more real-world crash cases than LCTGen. On 19 cases supported by AC3R, SAFE reproduces one additional crash case with statistically significant gains across five runs. It generates scenarios within 25 seconds and triggers violations after just 1 case (IDM) and 3 cases (PPO) in MetaDrive, as well as 1 case (Auto) in BeamNG.
• Software and its engineering → Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 31cb4518-f261-43a8-8f5c-5543d79a3220Builds on8
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee Manakul, Adian Liusie, Mark J. F. GalesEMNLP 2023 · 331 citations
- Testing of autonomous driving systems: where are we and where should we go?Guannan Lou, Yao Deng, Xi Zheng, Mengshi Zhang et al.FSE 2022 · 85 citations
- LawBreaker: An Approach for Specifying Traffic Laws and Fuzzing Autonomous VehiclesYang Sun, Christopher M. Poskitt, Jun Sun, Yuqi Chen et al.ASE 2022 · 46 citations
Related papers
- DiaVio: LLM-Empowered Diagnosis of Safety Violations in ADS Simulation TestingYou Lu, Yifan Tian, Yuyang Bi, Bihuan Chen et al.ISSTA 2024 · 9 citations
- SoVAR: Build Generalizable Scenarios from Accident Reports for Autonomous Driving TestingAn Guo, Yuan Zhou, Haoxiang Tian, Chunrong Fang et al.ASE 2024 · 12 citations
- SafeDriveRAG: Towards Safe Autonomous Driving with Knowledge Graph-based Retrieval-Augmented GenerationHao Ye, Mengshi Qi, Zhaohong Liu, Liang Liu et al.ACM MM 2025 · 6 citations
- Fixed-Point Guided ADS Scenario Generation via Multi-modal LLM Reasoning and Software TestingXudong Zhang, Shihao Zhu, Yan CaiISSTA 2026
- SafeAuto: Knowledge-Enhanced Safe Autonomous Driving with Multimodal Foundation ModelsJiawei Zhang, Xuan Yang, Taiqi Wang, Yu Yao et al.ICML 2025
