Towards Surveillance Video-and-Language Understanding: New Dataset, Baselines, and Challenges
Tongtong Yuan, Xuange Zhang, Kun Liu, Bo Liu, Chen Chen, Jian Jin, Zhenzhen Jiao
Abstract
Surveillance videos are important for public security. However, current surveillance video tasks mainly focus on classifying and localizing anomalous events. Existing methods are limited to detecting and classifying the predefined events with unsatisfactory semantic understanding, although they have obtained considerable performance. To address this issue, we propose a new research direction of surveillance video-and-language understanding (VALU), and construct the first multimodal surveillance video dataset. We manually annotate the real-world surveillance dataset UCF-Crime with fine-grained event content and timing. Our newly annotated dataset, UCA (UCF-Crime Annotation) 1 , contains 23,542 sentences, with an average length of 20 words, and its annotated videos are as long as 110.7 hours. Furthermore, we benchmark SOTA models for four multimodal tasks on this newly created dataset, which serve as new baselines for surveillance VALU. Through experiments, we find that mainstream models used in previously public datasets perform poorly on surveillance video, demonstrating new challenges in surveillance VALU. We also conducted experiments on multimodal anomaly detection. These results demonstrate that our multimodal surveillance learning can improve the performance of anomaly detection. All the experiments highlight the necessity of constructing this dataset to advance surveillance AI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 10a2e4e2-7259-4161-bab7-8290b90d3a02Cited by top-tier papers18
- HAWK: Learning to Understand Open-World Video AnomaliesJiaqi Tang, Hao Lu, Ruizheng Wu, Xiaogang Xu et al.NeurIPS 2024 · 71 citations
- Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-ThoughtChao Huang, Benfeng Wang, Wei Wang, Jie Wen et al.NeurIPS 2025 · 30 citations
- No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly DetectionZunkai Dai, Ke Li, Jiajia Liu, Jie Yang et al.CVPR 2026 · 6 citations
- VarCMP: Adapting Cross-Modal Pre-Training Models for Video Anomaly RetrievalPeng Wu, Wanshun Su, Xiangteng He, Peng Wang et al.AAAI 2025 · 6 citations
- Language-guided Open-world Video Anomaly Detection under Weak SupervisionZihao Liu, Xiaoyu Wu, Jianqin Wu, Xuxu Wang et al.ICLR 2026 · 5 citations
Builds on13
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
- VaTeX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearchXin Wang, Jiawei Wu, Jun-Kun Chen, Lei Li et al.ICCV 2019 · 688 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- SwinBERT: End-to-End Transformers with Sparse Attention for Video CaptioningKevin Lin, Linjie Li, Chung-Ching Lin, Faisal Ahmed et al.CVPR 2022 · 263 citations
- End-to-End Dense Video Captioning with Parallel DecodingTeng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng et al.ICCV 2021 · 238 citations
Related papers
- VALU: A Benchmark for Video Anomaly Temporal Localization and Understanding at Multiple Semantic LevelsYixiao He, Menghao Zhang, Haifeng Sun, Jing Wang et al.ACL 2026
- MGFN: Magnitude-Contrastive Glance-and-Focus Network for Weakly-Supervised Video Anomaly DetectionYingxian Chen, Zhengzhe Liu, Baoheng Zhang, Wilton W. T. Fok et al.AAAI 2023 · 221 citations
- Collaborative Learning of Anomalies with Privacy (CLAP) for Unsupervised Video Anomaly Detection: A New BaselineAnas Al-lahham, Muhammad Zaigham Zaheer, Nurbek Tastan, Karthik NandakumarCVPR 2024
- A New Comprehensive Benchmark for Semi-supervised Video Anomaly Detection and AnticipationCongqi Cao, Yue Lu, Peng Wang, Yanning ZhangCVPR 2023
- Harnessing Large Language Models for Training-Free Video Anomaly DetectionLuca Zanella, Willi Menapace, Massimiliano Mancini, Yiming Wang et al.CVPR 2024 · 57 citations
