Hermes: An Evidence-Driven Agentic Framework for Trustworthy and Explainable AI-Generated Video Detection
Shuaibo Li, Pengfei HAO, Hongtao Wu, Jianfeng Dong, Ping Li, Xiaohong Liu, Lei Zhu
摘要
Recent advances in generative video models have blurred the boundary between real and synthetic content, raising urgent concerns about digital authenticity. Multimodal large language models (MLLMs) are promising for AI-generated video (AIGV) detection due to their broad perceptual and reasoning capabilities; however, existing MLLM-based detectors remain prone to hallucinated evidence and unstable reasoning, leading to false alarms and generic, unverifiable explanations. To address these issues, we propose Hermes, an evidence-driven agentic framework for trustworthy and explainable AIGV detection. Hermes comprises three key components: (1) Adaptive Instance-Conditioned Detection Strategy Planning, (2) Evidence-Centric Reasoning and Verification, and (3) Graph-Grounded Evidence Deliberation. Specifically, Hermes analyzes each video and uses instance-conditioned retrieval-augmented generation to retrieve relevant forensic knowledge and compose a tailored detection strategy. It then constructs a verifiable Evidence Reasoning Graph (ERG) to keep the reasoning grounded in concrete video evidence. Finally, multi-agent deliberation audits and refines the ERG to reconcile conflicting evidence and improve the reliability of the final judgment. Together with a library of forensic tools, these components enable evidence-grounded authenticity judgments and structured, verifiable explanations. Extensive experiments show that Hermes, without task-specific training, achieves state-of-the-art performance and produces auditable explanations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- TF-ICON: Diffusion-Based Training-Free Cross-Domain Image CompositionShilin Lu, Yanzhu Liu, Adams Wai-Kin KongICCV 2023 · 被引用 214 次
- Spatiotemporal Inconsistency Learning for DeepFake Video DetectionZhihao Gu, Yang Chen, Taiping Yao, Shouhong Ding 等ACM MM 2021 · 被引用 175 次
- TALL: Thumbnail Layout for Deepfake Video DetectionYuting Xu, Jian Liang, Gengyun Jia, Ziming Yang 等ICCV 2023 · 被引用 133 次
相关 Paper
- Towards Explainable Fake Image Detection with Multi-Modal Large Language ModelsYikun Ji, Yan Hong, Jiahui Zhan, Haoxing Chen 等ACM MM 2025 · 被引用 3 次
- FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded ReasoningYikun Ji, Yan Hong, Qi Fan, Jun Lan 等ICLR 2026 · 被引用 9 次
- ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned RepresentationQing Huang, Zhipei Xu, Xuanyu Zhang, Xiangyu Yu 等CVPR 2026 · 被引用 3 次
- CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video DetectionHuidong Feng, Wentao Chen, Jie Chen, Xinqi Cai 等CVPR 2026 · 被引用 2 次
- Agentic Video Summarization via Self-Reflecting Multimodal UnderstandingMiaotian Guo, Shuguang Dou, Yin Li, Aidong Men 等CVPR 2026
