Deep Residual Injection for Full-Spectrum Forensic Signal Perception in Multimodal Large Language Models
Kaiqing Lin, Zhiyuan Yan, Ruoxin Chen, Ke-Yue Zhang, Yue Zhou, Caiyong Piao, Bin Li, Taiping Yao, Bo Wang, Youchang xiao, Shouhong Ding
摘要
Multimodal large language models (MLLMs) have been increasingly adopted in forensics for their robust semantic understanding. As AIgenerated images become realistic, semantic-level inconsistencies alone are often insufficient for reliable detection. This motivates a critical question: whether MLLMs can achieve full-spectrum forensic signal perception, i.e., capturing lowlevel generator artifacts without sacrificing pretrained semantic knowledge. We further perform a layer-wise analysis of forensic signal perception in MLLMs, showing that semantic information is primarily formed in the early-to-middle layers, whereas direct fine-tuning for artifact learning disrupts these semantic representations. Based on this insight, we propose Deep Visual Residual MLLM (Deep-VRM) to preserve early semantic processing while injecting artifact-specific visual signals as a residual path into an intermediate layer, where they are fused with semantic token representations and propagated through subsequent trainable layers. This enables later layers to jointly model semantic reasoning and signallevel forensic cues, and surprisingly, the model learns to adaptively leverage different levels of forensic signals depending on the input, achieving robust and generalizable detection performance. Extensive experiments show that our method achieves state-of-the-art across most benchmarks. The code and data are available at https:// github.com/KQL11/Deep-VRM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper30
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- DIRE for Diffusion-Generated Image DetectionZhendong Wang, Jianmin Bao, Wengang Zhou, Weilun Wang 等ICCV 2023 · 被引用 479 次
- Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Domain LearningChuangchuang Tan, Yao Zhao, Shikui Wei, Guanghua Gu 等AAAI 2024 · 被引用 232 次
- Rethinking the Up-Sampling Operations in CNN-Based Generative Network for Generalizable Deepfake DetectionChuangchuang Tan, Huan Liu, Yao Zhao, Shikui Wei 等CVPR 2024 · 被引用 126 次
相关 Paper
- Towards Explainable Fake Image Detection with Multi-Modal Large Language ModelsYikun Ji, Yan Hong, Jiahui Zhan, Haoxing Chen 等ACM MM 2025 · 被引用 3 次
- FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded ReasoningYikun Ji, Yan Hong, Qi Fan, Jun Lan 等ICLR 2026 · 被引用 9 次
- Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake DetectionPeipeng Yu, Jianwei Fei, Hui Gao, Xuan Feng 等ICML 2025
- Hermes: An Evidence-Driven Agentic Framework for Trustworthy and Explainable AI-Generated Video DetectionShuaibo Li, Pengfei HAO, Hongtao Wu, Jianfeng Dong 等ICML 2026
- ForgerySleuth: Empowering Multimodal Large Language Models for Image Manipulation DetectionZhihao Sun, Haoran Jiang, Haoran Chen, Yixin Cao 等NeurIPS 2025 · 被引用 16 次
