EchoTraffic: Enhancing Traffic Anomaly Understanding with Audio-Visual Insights
Zhenghao Xing, Hao Chen, Binzhu Xie, Jiaqi Xu, Ziyu Guo, Xuemiao Xu, Jianye Hao, Chi-Wing Fu, Xiaowei Hu, Pheng-Ann Heng
Abstract
Traffic Anomaly Understanding (TAU) is essential for improving public safety and transportation efficiency by enabling timely detection and response to incidents. Beyond existing methods, which rely largely on visual data, we propose to consider audio cues, a valuable source that offers strong hints to anomaly scenarios such as crashes and honking. Our contributions are twofold. First, we compile AV-TAU, the first large-scale audio-visual dataset for TAU, providing 29,865 traffic anomaly videos and 149,325 Q&A pairs, while supporting five essential TAU tasks. Second, we develop EchoTraffic, a multimodal LLM that integrates audio and visual data for TAU, through our audio-insight frame selector and dynamic connector to effectively extract crucial audio cues for anomaly understanding with a twophase training framework. Experimental results on AV-TAU manifest that EchoTraffic sets a new SOTA performance in TAU, outperforming the existing multimodal LLMs. Our contributions, including AV-TAU and EchoTraffic, pave a new direction for multimodal TAU.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly DetectionZunkai Dai, Ke Li, Jiajia Liu, Jie Yang et al.CVPR 2026 · 6 citations
- HAVE-Bench: Hierarchical Audio-Visual Evaluation from Perception to InteractionZhong Muyan, Erfei Cui, Sen Xing, Weiyun Wang et al.CVPR 2026
Builds on14
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li et al.ICLR 2024 · 467 citations
- Uncertainty-based Traffic Accident Anticipation with Spatio-Temporal Relational LearningWentao Bao, Qi Yu, Yu KongACM MM 2020 · 191 citations
Related papers
- TAU-106K: A New Dataset for Comprehensive Understanding of Traffic AccidentYixuan Zhou, Long Bai, Sijia Cai, Bing Deng et al.ICLR 2025
- Spotting the Unexpected (STU): A 3D LiDAR Dataset for Anomaly Segmentation in Autonomous DrivingAlexey Nekrasov, Malcolm Burdorf, Stewart Worrall, Bastian Leibe et al.CVPR 2025
- RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation SystemRunwei Guan, Rongsheng Hu, Shangshu Chen, Ningyuan Xiao et al.AAAI 2026 · 2 citations
- EventVAD: Training-Free Event-Aware Video Anomaly DetectionYihua Shao, Haojin He, Sijie Li, Siyu Chen et al.ACM MM 2025 · 19 citations
- Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any GranularityHuaxin Zhang, Xiaohao Xu, Xiang Wang, Jialong Zuo et al.CVPR 2025
