VILTA: A VLM-in-the-Loop Adversary for Enhancing Driving Policy Robustness
Qimao Chen, Fang Li, Shaoqing Xu, Zhiyi Lai, Zixun Xie, Yuechen Luo, Shengyin Jiang, Hanbing Li, Long Chen, Bing Wang, Yi Zhang, Zhi-Xin Yang
Abstract
The safe deployment of autonomous driving (AD) systems is fundamentally hindered by the long-tail problem, where rare yet critical driving scenarios are severely underrepresented in real-world data. Existing solutions including safety-critical scenario generation and closed-loop learning often rely on rule-based heuristics, resampling methods and generative models learned from offline datasets, limiting their ability to produce diverse and novel challenges. While recent works leverage Vision Language Models (VLMs) to produce scene descriptions that guide a separate, downstream model in generating hazardous trajectories for agents, such two-stage framework constrains the generative potential of VLMs, as the diversity of the final trajectories is ultimately limited by the generalization ceiling of the downstream algorithm. To overcome these limitations, we introduce VILTA (VLM-In-the-Loop Trajectory Adversary), a novel framework that integrates a VLM into the closed-loop training of AD agents. Unlike prior works, VILTA actively participates in the training loop by comprehending the dynamic driving environment and strategically generating challenging scenarios through direct, fine-grained editing of surrounding agents' future trajectories. This direct-editing approach fully leverages the VLM's powerful generalization capabilities to create a diverse curriculum of plausible yet challenging scenarios that extend beyond the scope of traditional methods. We demonstrate that our approach substantially enhances the safety and robustness of the resulting AD policy, particularly in its ability to navigate critical long-tail events.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c64733eb-adc2-45d3-a542-80334fc36fcaBuilds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsChenlin Meng, Yutong He, Yang Song, Jiaming Song et al.ICLR 2022 · 2,128 citations
- DenseTNT: End-to-end Trajectory Prediction from Dense Goal SetsJunru Gu, Chen Sun, Hang ZhaoICCV 2021 · 563 citations
- Evolving Curricula with Regret-Based Environment DesignJack Parker-Holder, Minqi Jiang, Michael Dennis, Mikayel Samvelyan et al.ICML 2022 · 175 citations
- ChatScene: Knowledge-Enabled Safety-Critical Scenario Generation for Autonomous VehiclesJiawei Zhang, Chejian Xu, Bo LiCVPR 2024 · 50 citations
Related papers
- CoC-VLA: Delving into Adversarial Domain Transfer for Explainable Autonomous Driving via Chain-of-Causality Visual-Language-Action ModelDapeng Zhang, Fei Shen, Rui Zhao, Yinda Chen et al.NeurIPS 2025 · 8 citations
- The Blind Spot of Adaptation: Quantifying and Mitigating Forgetting in Fine-tuned Driving ModelsRunhao Mao, Hanshi Wang, Yixiang Yang, Qianli Ma et al.CVPR 2026 · 1 citation
- Driving with Advice: Large Model as Motion Advisor for Joint PlanningJunyin Wang, Jinlei Yu, Hao Lin, Huikai Liu et al.AAAI 2026
- SGDrive: Scene-to-Goal Hierarchical World Cognition for Autonomous Drivingjingyu li, Junjie Wu, Dongnan Hu, Xiangkai Huang et al.CVPR 2026 · 36 citations
- Generating Traffic Scenarios via In-Context Learning to Learn Better Motion PlannerAizierjiang AiersilanAAAI 2025 · 6 citations
