SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations
Zhaorun Chen, Francesco Pinto, Minzhou Pan, Bo Li
Abstract
With the rise of generative AI and rapid growth of high-quality video generation, video guardrails have become more crucial than ever to ensure safety and security across platforms. Current video guardrails, however, are either overly simplistic, relying on pure classification models trained on simple policies with limited unsafe categories, which lack detailed explanations, or prompting multimodal large language models (MLLMs) with long safety guidelines, which are inefficient and impractical for guardrailing real-world content. To bridge this gap, we propose SAFEWATCH, an efficient MLLM-based video guardrail model designed to follow customized safety policies and provide multi-label video guardrail outputs with content-specific explanations in a zero-shot manner. In particular, unlike traditional MLLM-based guardrails that encode all safety policies autoregressively, causing inefficiency and bias, SAFEWATCH uniquely encodes each policy chunk in parallel and eliminates their position bias such that all policies are attended simultaneously with equal importance. In addition, to improve efficiency and accuracy, SAFEWATCH incorporates a policy-aware visual token pruning algorithm that adaptively selects the most relevant video tokens for each policy, discarding noisy or irrelevant information. This allows for more focused, policy-compliant guardrail with significantly reduced computational overhead. Considering the limitations of existing video guardrail benchmarks, we propose SAFEWATCH-BENCH, a large-scale video guardrail benchmark comprising over 2M videos spanning six safety categories which covers over 30 tasks to ensure a comprehensive coverage of all potential safety scenarios. We have conducted extensive experiments, showing that SAFEWATCH outperforms all SOTA video guardrails on SAFEWATCH-BENCH by 28.2%, and achieves a 13.6% improvement on existing benchmarks, all while reducing inference costs by an average of 10%. SAFE-WATCH also demonstrates strong policy-following abilities and outperforms previous SOTAs by 5.6% and 15.6% in zero-shot generalizability to new policies and new prompting tasks. Additionally, both LLM-as-a-judge and human evaluators confirm the high quality of the explanations provided by SAFEWATCH. Our project is open-sourced at https://safewatch-aiguard.github.io .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMsZichen Wen, Jiashu Qu, Zhaorun Chen, Xiaoya Lu et al.ICLR 2026 · 32 citations
- Efficient Multi-modal Large Language Models via Progressive Consistency DistillationZichen Wen, Shaobo Wang, Yufa Zhou, Junyuan Zhang et al.NeurIPS 2025 · 26 citations
- VidGuard-R1: AI-Generated Video Detection and Explanation via Reasoning MLLMs and RLKyoungjun Park, Yifan Yang, Juheon Yi, Shicheng Zheng et al.ICLR 2026 · 16 citations
- Towards Policy-Adaptive Image Guardrail: Benchmark and MethodCaiyong Piao, Zhiyuan Yan, Haoming Xu, Yunzhen Zhao et al.CVPR 2026 · 7 citations
- ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play AttacksZhaorun Chen, Xun Liu, Mintong Kang, Jiawei Zhang et al.ICLR 2026 · 5 citations
Builds on17
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo et al.NeurIPS 2024 · 1,004 citations
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge BasesZhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song et al.NeurIPS 2024 · 539 citations
- Make-A-Video: Text-to-Video Generation without Text-Video DataUriel Singer, Adam Polyak, Thomas Hayes, Xi Yin et al.ICLR 2023 · 313 citations
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai et al.EMNLP 2022 · 239 citations
Related papers
- SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore MechanismBeitao Chen, Xinyu Lyu, Shengming Yuan, Jingkuan Song et al.NeurIPS 2025 · 14 citations
- Breaking Multimodal LLM Safety via Video-Driven PromptingDong Wang, XIANGYU HE, Xinqi Lyu, Bin XiaoCVPR 2026
- GuardBench: A Large-Scale Benchmark for Guardrail ModelsElias Bassani, Ignacio SanchezEMNLP 2024 · 7 citations
- From Evaluation to Defense: Advancing Safety in Video Large Language ModelsYiwei Sun, Peiqi Jiang, Chuanbin Liu, Luohao Lin et al.ICLR 2026 · 2 citations
- MLLM-as-a-Judge for Image Safety without Human LabelingZhenting Wang, Shuming Hu, Shiyu Zhao, Xiaowen Lin et al.CVPR 2025
