SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations
Zhaorun Chen, Francesco Pinto, Minzhou Pan, Bo Li
摘要
With the rise of generative AI and rapid growth of high-quality video generation, video guardrails have become more crucial than ever to ensure safety and security across platforms. Current video guardrails, however, are either overly simplistic, relying on pure classification models trained on simple policies with limited unsafe categories, which lack detailed explanations, or prompting multimodal large language models (MLLMs) with long safety guidelines, which are inefficient and impractical for guardrailing real-world content. To bridge this gap, we propose SAFEWATCH, an efficient MLLM-based video guardrail model designed to follow customized safety policies and provide multi-label video guardrail outputs with content-specific explanations in a zero-shot manner. In particular, unlike traditional MLLM-based guardrails that encode all safety policies autoregressively, causing inefficiency and bias, SAFEWATCH uniquely encodes each policy chunk in parallel and eliminates their position bias such that all policies are attended simultaneously with equal importance. In addition, to improve efficiency and accuracy, SAFEWATCH incorporates a policy-aware visual token pruning algorithm that adaptively selects the most relevant video tokens for each policy, discarding noisy or irrelevant information. This allows for more focused, policy-compliant guardrail with significantly reduced computational overhead. Considering the limitations of existing video guardrail benchmarks, we propose SAFEWATCH-BENCH, a large-scale video guardrail benchmark comprising over 2M videos spanning six safety categories which covers over 30 tasks to ensure a comprehensive coverage of all potential safety scenarios. We have conducted extensive experiments, showing that SAFEWATCH outperforms all SOTA video guardrails on SAFEWATCH-BENCH by 28.2%, and achieves a 13.6% improvement on existing benchmarks, all while reducing inference costs by an average of 10%. SAFE-WATCH also demonstrates strong policy-following abilities and outperforms previous SOTAs by 5.6% and 15.6% in zero-shot generalizability to new policies and new prompting tasks. Additionally, both LLM-as-a-judge and human evaluators confirm the high quality of the explanations provided by SAFEWATCH. Our project is open-sourced at https://safewatch-aiguard.github.io .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMsZichen Wen, Jiashu Qu, Zhaorun Chen, Xiaoya Lu 等ICLR 2026 · 被引用 32 次
- Efficient Multi-modal Large Language Models via Progressive Consistency DistillationZichen Wen, Shaobo Wang, Yufa Zhou, Junyuan Zhang 等NeurIPS 2025 · 被引用 26 次
- VidGuard-R1: AI-Generated Video Detection and Explanation via Reasoning MLLMs and RLKyoungjun Park, Yifan Yang, Juheon Yi, Shicheng Zheng 等ICLR 2026 · 被引用 16 次
- Towards Policy-Adaptive Image Guardrail: Benchmark and MethodCaiyong Piao, Zhiyuan Yan, Haoming Xu, Yunzhen Zhao 等CVPR 2026 · 被引用 7 次
- ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play AttacksZhaorun Chen, Xun Liu, Mintong Kang, Jiawei Zhang 等ICLR 2026 · 被引用 5 次
它引用的顶会 Paper17
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo 等NeurIPS 2024 · 被引用 1,004 次
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge BasesZhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song 等NeurIPS 2024 · 被引用 539 次
- Make-A-Video: Text-to-Video Generation without Text-Video DataUriel Singer, Adam Polyak, Thomas Hayes, Xi Yin 等ICLR 2023 · 被引用 313 次
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai 等EMNLP 2022 · 被引用 239 次
相关 Paper
- SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore MechanismBeitao Chen, Xinyu Lyu, Shengming Yuan, Jingkuan Song 等NeurIPS 2025 · 被引用 14 次
- Breaking Multimodal LLM Safety via Video-Driven PromptingDong Wang, XIANGYU HE, Xinqi Lyu, Bin XiaoCVPR 2026
- GuardBench: A Large-Scale Benchmark for Guardrail ModelsElias Bassani, Ignacio SanchezEMNLP 2024 · 被引用 7 次
- From Evaluation to Defense: Advancing Safety in Video Large Language ModelsYiwei Sun, Peiqi Jiang, Chuanbin Liu, Luohao Lin 等ICLR 2026 · 被引用 2 次
- MLLM-as-a-Judge for Image Safety without Human LabelingZhenting Wang, Shuming Hu, Shiyu Zhao, Xiaowen Lin 等CVPR 2025
