VSG-Safe: Spotting NSFW Video through Cross-Frame Evidence
Yuyang Zhang, Xudong Jiang, Yuxuan Song, Yuxiang Sun, Yihao Huang, Run Wang, Shundi Xiao, Lina Wang
摘要
Recent advances in text-to-video (T2V) models enable highfidelity videos that closely follow textual prompts. However, this expands practical applications while amplifying serious security and societal concerns from the automated synthesis of visual content that may be inappropriate in certain usage contexts, such as public or workplace settings, including sexual or violent content (e.g., the Grok can generate sexual videos in the "Spicy" mode). We observe that such visual content is often distributed across frames, embedded in visual entities, their attributes, and inter-entity relations. In contrast, existing moderation pipelines primarily treat video content as either individual frames or raw frame sequences, overlooking the fact that critical semantics can manifest through the combination of specific frames. This gap prevents them from reasoning across frames, confining detection to low-level visual cues, such as gore or explicit conflict, and causing frequent failures when cross-frame inference is required, including illegal activities or threats. To address these limitations, we propose leveraging scene graphs as the core intermediate semantic representation. Scene graphs naturally encode entities, their attributes, and inter-entity relationships, while also supporting reasoning over cross-frame content. Grounded on this insight, we further propose VSG-Safe, a novel scene-graph-driven framework for T2V content moderation. Concretely, our approach first extracts cross-frame content from videos to build scene graphs. With these graphs, we leverage a graph-oriented model to jointly capture entities, attributes, and inter-entity relations, enabling effective detection. To evaluate its effectiveness, we conduct extensive experiments on both SOTA benchmarks and our self-constructed video datasets. VSG-Safe attains an average F1-score of 97.62%, outperforming seven baselines by 42.32% on average. Disclaimer: This paper contains visual content that might be offensive to some readers, such as sexual and violent content. Although we censor and mask Not-Safe-for-Work (NSFW) imagery, reader discretion is advised.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- How Attentive are Graph Attention Networks?Shaked Brody, Uri Alon, Eran YahavICLR 2022 · 被引用 1,717 次
- MotionBooth: Motion-Aware Customized Text-to-Video GenerationJianzong Wu, Xiangtai Li, Yanhong Zeng, Jiangning Zhang 等NeurIPS 2024 · 被引用 114 次
- Do Different Tracking Tasks Require Different Appearance Models?Zhongdao Wang, Hengshuang Zhao, Ya-Li Li, Shengjin Wang 等NeurIPS 2021 · 被引用 107 次
相关 Paper
- Bringing Real-World Relations into Video Generation with Graph-Structured KnowledgeJoonhyung Park, Jaeyun Song, Sihwan Park, Eunho YangACL 2026
- Target Adaptive Context Aggregation for Video Scene Graph GenerationYao Teng, Limin Wang, Zhifeng Li, Gangshan WuICCV 2021 · 被引用 80 次
- TEAR: Temporal-aware Automated Red-teaming for Text-to-Video ModelsJiaming He, Guanyu Hou, Hongwei Li, Zhicong Huang 等CVPR 2026 · 被引用 3 次
- (2.5+1)D Spatio-Temporal Scene Graphs for Video Question AnsweringAnoop Cherian, Chiori Hori, Tim K. Marks, Jonathan Le RouxAAAI 2022 · 被引用 48 次
- From Evaluation to Defense: Advancing Safety in Video Large Language ModelsYiwei Sun, Peiqi Jiang, Chuanbin Liu, Luohao Lin 等ICLR 2026 · 被引用 2 次
