Abductive Ego-View Accident Video Understanding for Safe Driving Perception
Jianwu Fang, Lei-Lei Li, Junfei Zhou, Junbin Xiao, Hongkai Yu, Chen Lv, Jianru Xue, Tat-Seng Chua
摘要
We present MM-AU, a novel dataset for Multi-Modal Accident video Understanding. MM-AU contains 11,727 in-the-wild ego-view accident videos, each with temporally aligned text descriptions. We annotate over 2.23 million object boxes and 58,650 pairs of video-based accident reasons, covering 58 accident categories. MM-AU supports various accident understanding tasks, particularly multimodal video diffusion to understand accident causeeffect chains for safe driving. With MM-AU, we present an Abductive accident Video understanding framework for Safe Driving perception (AdVersa-SD). AdVersa-SD performs video diffusion via an Object-Centric Video Diffusion (OAVD) method which is driven by an abductive CLIP model. This model involves a contrastive interaction loss to learn the pair co-occurrence of normal, near-accident, accident frames with the corresponding text descriptions, such as accident reasons, prevention advice, and accident categories. OAVD enforces the object region learning while fixing the content of the original frame background in video generation, to find the dominant objects for certain accidents. Extensive experiments verify the abductive ability of AdVersa-SD and the superiority of OAVD against the stateof-the-art diffusion models. Additionally, we provide careful benchmark evaluations for object detection and accident reason answering since AdVersa-SD relies on precise object and accident reason information.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- StreamForest: Efficient Online Video Understanding with Persistent Event MemoryXiangyu Zeng, Kefan Qiu, Qingyu Zhang, Xinhao Li 等NeurIPS 2025 · 被引用 79 次
- FineSports: A Multi-Person Hierarchical Sports Video Dataset for Fine-Grained Action UnderstandingJinglin Xu, Guohao Zhao, Sibo Yin, Wenhao Zhou 等CVPR 2024 · 被引用 12 次
- Accident Anticipation via Temporal Occurrence PredictionTianhao Zhao, Yiyang Zou, Zihao Mao, Peilun Xiao 等NeurIPS 2025 · 被引用 6 次
- Causal-Entity Reflected Egocentric Traffic Accident Video SynthesisLei-Lei Li, Jianwu Fang, Junbin Xiao, Shanmin Pang 等ICCV 2025 · 被引用 4 次
- RiskProp: Collision-Anchored Self-Supervised Risk Propagation For Early Accident AnticipationYiyang Zou, Tianhao Zhao, Peilun Xiao, Hongyu Jin 等CVPR 2026 · 被引用 4 次
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
相关 Paper
- TAU-106K: A New Dataset for Comprehensive Understanding of Traffic AccidentYixuan Zhou, Long Bai, Sijia Cai, Bing Deng 等ICLR 2025
- On Learning Multi-Modal Forgery Representation for Diffusion Generated Video DetectionXiufeng Song, Xiao Guo, Jiache Zhang, Qirui Li 等NeurIPS 2024 · 被引用 63 次
- Towards Safer and Understandable Driver Intention PredictionMukilan Karuppasamy, Shankar Gangisetty, Shyam Nandan Rai, Carlo Masone 等ICCV 2025 · 被引用 2 次
- CRASH: Crash Recognition and Anticipation System Harnessing with Context-Aware and Temporal Focus AttentionsHaicheng Liao, Haoyu Sun, Huanming Shen, Chengyue Wang 等ACM MM 2024 · 被引用 10 次
- World Knowledge-Enhanced Reasoning Using Instruction-Guided Interactor in Autonomous DrivingMingliang Zhai, Cheng Li, Zengyuan Guo, Ningrui Yang 等AAAI 2025 · 被引用 9 次
