Hermes: An Evidence-Driven Agentic Framework for Trustworthy and Explainable AI-Generated Video Detection
Shuaibo Li, Pengfei HAO, Hongtao Wu, Jianfeng Dong, Ping Li, Xiaohong Liu, Lei Zhu
Abstract
Recent advances in generative video models have blurred the boundary between real and synthetic content, raising urgent concerns about digital authenticity. Multimodal large language models (MLLMs) are promising for AI-generated video (AIGV) detection due to their broad perceptual and reasoning capabilities; however, existing MLLM-based detectors remain prone to hallucinated evidence and unstable reasoning, leading to false alarms and generic, unverifiable explanations. To address these issues, we propose Hermes, an evidence-driven agentic framework for trustworthy and explainable AIGV detection. Hermes comprises three key components: (1) Adaptive Instance-Conditioned Detection Strategy Planning, (2) Evidence-Centric Reasoning and Verification, and (3) Graph-Grounded Evidence Deliberation. Specifically, Hermes analyzes each video and uses instance-conditioned retrieval-augmented generation to retrieve relevant forensic knowledge and compose a tailored detection strategy. It then constructs a verifiable Evidence Reasoning Graph (ERG) to keep the reasoning grounded in concrete video evidence. Finally, multi-agent deliberation audits and refines the ERG to reconcile conflicting evidence and improve the reliability of the final judgment. Together with a library of forensic tools, these components enable evidence-grounded authenticity judgments and structured, verifiable explanations. Extensive experiments show that Hermes, without task-specific training, achieves state-of-the-art performance and produces auditable explanations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 283d9df1-684d-49d8-83a2-67d06c12ba58Builds on19
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- TF-ICON: Diffusion-Based Training-Free Cross-Domain Image CompositionShilin Lu, Yanzhu Liu, Adams Wai-Kin KongICCV 2023 · 214 citations
- Spatiotemporal Inconsistency Learning for DeepFake Video DetectionZhihao Gu, Yang Chen, Taiping Yao, Shouhong Ding et al.ACM MM 2021 · 175 citations
- TALL: Thumbnail Layout for Deepfake Video DetectionYuting Xu, Jian Liang, Gengyun Jia, Ziming Yang et al.ICCV 2023 · 133 citations
Related papers
- Towards Explainable Fake Image Detection with Multi-Modal Large Language ModelsYikun Ji, Yan Hong, Jiahui Zhan, Haoxing Chen et al.ACM MM 2025 · 3 citations
- FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded ReasoningYikun Ji, Yan Hong, Qi Fan, Jun Lan et al.ICLR 2026 · 9 citations
- ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned RepresentationQing Huang, Zhipei Xu, Xuanyu Zhang, Xiangyu Yu et al.CVPR 2026 · 3 citations
- CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video DetectionHuidong Feng, Wentao Chen, Jie Chen, Xinqi Cai et al.CVPR 2026 · 2 citations
- Agentic Video Summarization via Self-Reflecting Multimodal UnderstandingMiaotian Guo, Shuguang Dou, Yin Li, Aidong Men et al.CVPR 2026
