SafeAuto: Knowledge-Enhanced Safe Autonomous Driving with Multimodal Foundation Models
Jiawei Zhang, Xuan Yang, Taiqi Wang, Yu Yao, Aleksandr Petiushko, Bo Li
摘要
Traditional autonomous driving systems often struggle to connect high-level reasoning with lowlevel control, leading to suboptimal and sometimes unsafe behaviors. Recent advances in multimodal large language models (MLLMs), which process both visual and textual data, offer an opportunity to unify perception and reasoning. However, effectively embedding precise safety knowledge into MLLMs for autonomous driving remains a significant challenge. To address this, we propose SafeAuto, a framework that enhances MLLM-based autonomous driving by incorporating both unstructured and structured knowledge. First, we introduce a Position-Dependent Cross-Entropy (PDCE) loss to improve low-level control signal predictions when values are represented as text. Second, to explicitly integrate safety knowledge, we develop a reasoning component that translates traffic rules into first-order logic (e.g., "red light =⇒ stop") and embeds them into a probabilistic graphical model (e.g., Markov Logic Network) to verify predicted actions using recognized environmental attributes. Additionally, our Multimodal Retrieval-Augmented Generation (RAG) model leverages video, control signals, and environmental attributes to learn from past driving experiences. Integrating PDCE, MLN, and Multimodal RAG, SafeAuto outperforms existing baselines across multiple datasets, enabling more accurate, reliable, and safer autonomous driving. The code is available at https:// github.com/AI-secure/SafeAuto .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous DrivingYingyan Li, Shuyao Shang, Weisong Liu, Bing Zhan 等ICLR 2026 · 被引用 134 次
- The Blind Spot of Adaptation: Quantifying and Mitigating Forgetting in Fine-tuned Driving ModelsRunhao Mao, Hanshi Wang, Yixiang Yang, Qianli Ma 等CVPR 2026 · 被引用 1 次
- Bird's-eye-view Informed Reasoning DriverYinuo Wang, Mining Tan, Yuanxin Zhong, Wang zhitao 等ICLR 2026
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
相关 Paper
- SafeDriveRAG: Towards Safe Autonomous Driving with Knowledge Graph-based Retrieval-Augmented GenerationHao Ye, Mengshi Qi, Zhaohong Liu, Liang Liu 等ACM MM 2025 · 被引用 6 次
- DriveCombo: Benchmarking Compositional Traffic Rule Reasoning in Autonomous DrivingEnhui Ma, Jiahuan Zhang, Guantian Zheng, Tao Tang 等CVPR 2026 · 被引用 2 次
- Reasoning-Driven Anomaly Detection and Localization with Image-Level SupervisionYizhou Jin, Yuezhu Feng, Jinjin Zhang, Peng Wang 等CVPR 2026 · 被引用 4 次
- SAFE: Harnessing LLM for Scenario-Driven ADS Testing from Multimodal Crash DataSiwei Luo, Yang Zhang, Yao Deng, Linfeng Liang 等ICSE 2026
- World Knowledge-Enhanced Reasoning Using Instruction-Guided Interactor in Autonomous DrivingMingliang Zhai, Cheng Li, Zengyuan Guo, Ningrui Yang 等AAAI 2025 · 被引用 9 次
