USENIX Security2026Top-tier venue
VSG-Safe: Spotting NSFW Video through Cross-Frame Evidence
Yuyang Zhang, Xudong Jiang, Yuxuan Song, Yuxiang Sun, Yihao Huang, Run Wang, Shundi Xiao, Lina Wang
Abstract
Recent advances in text-to-video (T2V) models enable highfidelity videos that closely follow textual prompts. However, this expands practical applications while amplifying serious security and societal concerns from the automated synthesis of visual content that may be inappropriate in certain usage contexts, such as public or workplace settings, including sexual or violent content (e.g., the Grok can generate sexual videos in the "Spicy" mode). We observe that such visual content is often distributed across frames, embedded in visual entities, their attributes, and inter-entity relations. In contrast, existing moderation pipelines primarily treat video content as either individual frames or raw frame sequences, overlooking the fact that critical semantics can manifest through the combination of specific frames. This gap prevents them from reasoning across frames, confining detection to low-level visual cues, such as gore or explicit conflict, and causing frequent failures when cross-frame inference is required, including illegal activities or threats. To address these limitations, we propose leveraging scene graphs as the core intermediate semantic representation. Scene graphs naturally encode entities, their attributes, and inter-entity relationships, while also supporting reasoning over cross-frame content. Grounded on this insight, we further propose VSG-Safe, a novel scene-graph-driven framework for T2V content moderation. Concretely, our approach first extracts cross-frame content from videos to build scene graphs. With these graphs, we leverage a graph-oriented model to jointly capture entities, attributes, and inter-entity relations, enabling effective detection. To evaluate its effectiveness, we conduct extensive experiments on both SOTA benchmarks and our self-constructed video datasets. VSG-Safe attains an average F1-score of 97.62%, outperforming seven baselines by 42.32% on average. Disclaimer: This paper contains visual content that might be offensive to some readers, such as sexual and violent content. Although we censor and mask Not-Safe-for-Work (NSFW) imagery, reader discretion is advised.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 33ba0c25-d381-49e3-ba9b-087184bdc9bbBuilds on17
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- How Attentive are Graph Attention Networks?Shaked Brody, Uri Alon, Eran YahavICLR 2022 · 1,717 citations
- MotionBooth: Motion-Aware Customized Text-to-Video GenerationJianzong Wu, Xiangtai Li, Yanhong Zeng, Jiangning Zhang et al.NeurIPS 2024 · 114 citations
- Do Different Tracking Tasks Require Different Appearance Models?Zhongdao Wang, Hengshuang Zhao, Ya-Li Li, Shengjin Wang et al.NeurIPS 2021 · 107 citations
Related papers
- Bringing Real-World Relations into Video Generation with Graph-Structured KnowledgeJoonhyung Park, Jaeyun Song, Sihwan Park, Eunho YangACL 2026
- Target Adaptive Context Aggregation for Video Scene Graph GenerationYao Teng, Limin Wang, Zhifeng Li, Gangshan WuICCV 2021 · 80 citations
- TEAR: Temporal-aware Automated Red-teaming for Text-to-Video ModelsJiaming He, Guanyu Hou, Hongwei Li, Zhicong Huang et al.CVPR 2026 · 3 citations
- (2.5+1)D Spatio-Temporal Scene Graphs for Video Question AnsweringAnoop Cherian, Chiori Hori, Tim K. Marks, Jonathan Le RouxAAAI 2022 · 48 citations
- From Evaluation to Defense: Advancing Safety in Video Large Language ModelsYiwei Sun, Peiqi Jiang, Chuanbin Liu, Luohao Lin et al.ICLR 2026 · 2 citations
