Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
Yifei Li, Wenzhao Zheng, Yanran Zhang, Runze Sun, Yu Zheng, Lei Chen, Jie Zhou, Jiwen Lu
摘要
The misuse of AI-driven video generation technologies has raised serious social concerns, highlighting the urgent need for reliable AI-generated video detectors. However, most existing methods are limited to binary classification and lack the necessary explanations for human interpretation. In this paper, we present , a specialized multimodal large language model (MLLM) that identifies human-perceivable visual artifacts in AI-generated videos and leverages them as grounded evidence for both detection and explanation. To support this objective, we construct for Supervised Fine-Tuning (SFT), which represents the first large-scale AI-generated video artifact dataset with fine-grained human annotations.We then develop a two-stage training strategy that systematically enhances our model's spatio-temporal artifact perception, explanation capability, and detection accuracy. To comprehensively evaluate Skyra, we introduce , a benchmark comprising 3K high-quality samples generated by over ten state-of-the-art video generators. Extensive experiments demonstrate that Skyra surpasses existing methods across multiple benchmarks, while our evaluation yields valuable insights for advancing explainable AI-generated video detection. Our code, models, and datasets will be made publicly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image DetectionYanran Zhang, Wenzhao Zheng, Yifei Li, Bingyao Yu 等CVPR 2026 · 被引用 3 次
- Hermes: An Evidence-Driven Agentic Framework for Trustworthy and Explainable AI-Generated Video DetectionShuaibo Li, Pengfei HAO, Hongtao Wu, Jianfeng Dong 等ICML 2026
- Real Data Lies: Unveiling and Closing the Quality Shortcut in Generalizable AI-Generated Video DetectionZiyuan Fang, Tianyi Wei, Guanjie Wang, Weiming Zhang 等ICML 2026
它引用的顶会 Paper37
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
相关 Paper
- FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded ReasoningYikun Ji, Yan Hong, Qi Fan, Jun Lan 等ICLR 2026 · 被引用 9 次
- Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact ExplanationSiwei Wen, Junyan Ye, Peilin Feng, Hengrui Kang 等NeurIPS 2025 · 被引用 82 次
- CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video DetectionHuidong Feng, Wentao Chen, Jie Chen, Xinqi Cai 等CVPR 2026 · 被引用 2 次
- LEGION: Learning to Ground and Explain for Synthetic Image DetectionHengrui Kang, Siwei Wen, Zichen Wen, Junyan Ye 等ICCV 2025 · 被引用 3 次
- GenVidBench: A 6-Million Benchmark for AI-Generated Video DetectionZhenliang Ni, Qiangyu Yan, Mouxiao Huang, Tianning Yuan 等AAAI 2026 · 被引用 13 次
