Towards Surveillance Video-and-Language Understanding: New Dataset, Baselines, and Challenges
Tongtong Yuan, Xuange Zhang, Kun Liu, Bo Liu, Chen Chen, Jian Jin, Zhenzhen Jiao
摘要
Surveillance videos are important for public security. However, current surveillance video tasks mainly focus on classifying and localizing anomalous events. Existing methods are limited to detecting and classifying the predefined events with unsatisfactory semantic understanding, although they have obtained considerable performance. To address this issue, we propose a new research direction of surveillance video-and-language understanding (VALU), and construct the first multimodal surveillance video dataset. We manually annotate the real-world surveillance dataset UCF-Crime with fine-grained event content and timing. Our newly annotated dataset, UCA (UCF-Crime Annotation) 1 , contains 23,542 sentences, with an average length of 20 words, and its annotated videos are as long as 110.7 hours. Furthermore, we benchmark SOTA models for four multimodal tasks on this newly created dataset, which serve as new baselines for surveillance VALU. Through experiments, we find that mainstream models used in previously public datasets perform poorly on surveillance video, demonstrating new challenges in surveillance VALU. We also conducted experiments on multimodal anomaly detection. These results demonstrate that our multimodal surveillance learning can improve the performance of anomaly detection. All the experiments highlight the necessity of constructing this dataset to advance surveillance AI.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- HAWK: Learning to Understand Open-World Video AnomaliesJiaqi Tang, Hao Lu, Ruizheng Wu, Xiaogang Xu 等NeurIPS 2024 · 被引用 71 次
- Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-ThoughtChao Huang, Benfeng Wang, Wei Wang, Jie Wen 等NeurIPS 2025 · 被引用 30 次
- No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly DetectionZunkai Dai, Ke Li, Jiajia Liu, Jie Yang 等CVPR 2026 · 被引用 6 次
- VarCMP: Adapting Cross-Modal Pre-Training Models for Video Anomaly RetrievalPeng Wu, Wanshun Su, Xiangteng He, Peng Wang 等AAAI 2025 · 被引用 6 次
- Language-guided Open-world Video Anomaly Detection under Weak SupervisionZihao Liu, Xiaoyu Wu, Jianqin Wu, Xuxu Wang 等ICLR 2026 · 被引用 5 次
它引用的顶会 Paper13
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei 等CVPR 2022 · 被引用 1,847 次
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li 等ICCV 2019 · 被引用 688 次
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 被引用 579 次
- SwinBERT: End-to-End Transformers with Sparse Attention for Video CaptioningKevin Lin, Linjie Li, Chung-Ching Lin, Faisal Ahmed 等CVPR 2022 · 被引用 263 次
- End-to-End Dense Video Captioning with Parallel DecodingTeng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng 等ICCV 2021 · 被引用 238 次
相关 Paper
- VALU: A Benchmark for Video Anomaly Temporal Localization and Understanding at Multiple Semantic LevelsYixiao He, Menghao Zhang, Haifeng Sun, Jing Wang 等ACL 2026
- MGFN: Magnitude-Contrastive Glance-and-Focus Network for Weakly-Supervised Video Anomaly DetectionYingxian Chen, Zhengzhe Liu, Baoheng Zhang, Wilton W. T. Fok 等AAAI 2023 · 被引用 221 次
- Collaborative Learning of Anomalies with Privacy (CLAP) for Unsupervised Video Anomaly Detection: A New BaselineAnas Al-lahham, Muhammad Zaigham Zaheer, Nurbek Tastan, Karthik NandakumarCVPR 2024
- A New Comprehensive Benchmark for Semi-supervised Video Anomaly Detection and AnticipationCongqi Cao, Yue Lu, Peng Wang, Yanning ZhangCVPR 2023
- Harnessing Large Language Models for Training-Free Video Anomaly DetectionLuca Zanella, Willi Menapace, Massimiliano Mancini, Yiming Wang 等CVPR 2024 · 被引用 57 次
