Detecting and Grounding Multi-Modal Media Manipulation
Rui Shao, Tianxing Wu, Ziwei Liu
摘要
Misinformation has become a pressing issue. Fake media, in both visual and textual forms, is widespread on the web. While various deepfake detection and text fake news detection methods have been proposed, they are only designed for single-modality forgery based on binary classification, let alone analyzing and reasoning subtle forgery traces across different modalities. In this paper, we highlight a new research problem for multi-modal fake media, namely Detecting and Grounding Multi-Modal Media Manipulation (DGM 4 ). DGM 4 aims to not only detect the authenticity of multi-modal media, but also ground the manipulated content (i.e., image bounding boxes and text tokens), which requires deeper reasoning of multi-modal media manipulation. To support a large-scale investigation, we construct the first DGM 4 dataset, where image-text pairs are manipulated by various approaches, with rich anno-* This work was done at S-Lab, Nanyang Technological University † Corresponding author tation of diverse manipulations. Moreover, we propose a novel HierArchical Multi-modal Manipulation rEasoning tRansformer (HAMMER) to fully capture the fine-grained interaction between different modalities. HAMMER performs 1) manipulation-aware contrastive learning between two uni-modal encoders as shallow manipulation reasoning, and 2) modality-aware cross-attention by multi-modal aggregator as deep manipulation reasoning. Dedicated manipulation detection and grounding heads are integrated from shallow to deep levels based on the interacted multimodal information. Finally, we build an extensive benchmark and set up rigorous evaluation metrics for this new research problem. Comprehensive experiments demonstrate the superiority of our model; several valuable observations are also revealed to facilitate future research in multi-modal media manipulation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper46
- Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon TasksZaijing Li, Yuquan Xie, Rui Shao, Gongwei Chen 等NeurIPS 2024 · 被引用 104 次
- Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact ExplanationSiwei Wen, Junyan Ye, Peilin Feng, Hengrui Kang 等NeurIPS 2025 · 被引用 82 次
- DiffusionFake: Enhancing Generalization in Deepfake Detection via Guided Stable DiffusionKe Sun, Shen Chen, Taiping Yao, Hong Liu 等NeurIPS 2024 · 被引用 57 次
- MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language ModelsLeyang Shen, Gongwei Chen, Rui Shao, Weili Guan 等NeurIPS 2024 · 被引用 55 次
- Sniffer: Multimodal Large Language Model for Explainable Out-of-Context Misinformation DetectionPeng Qi, Zehong Yan, Wynne Hsu, Mong-Li LeeCVPR 2024 · 被引用 54 次
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess 等ICCV 2019 · 被引用 2,966 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or 等ICCV 2021 · 被引用 1,437 次
相关 Paper
- Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media ManipulationYiheng Li, Yang Yang, Zichang Tan, Huan Liu 等CVPR 2025
- ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and GroundingZhenxing Zhang, Yaxiong Wang, Lechao Cheng, Zhun Zhong 等CVPR 2025
- Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal ManipulationsJinjie Shen, Yaxiong Wang, Lechao Cheng, Nan Pu 等ACM MM 2025 · 被引用 2 次
- Cultivating Forensic Reasoning for Generalizable Multimodal Manipulation DetectionYuchen Zhang, Yaxiong Wang, Kecheng Han, Yujiao Wu 等ACL 2026 · 被引用 1 次
- CORE: Conflict-Oriented Reasoning for General Multimodal Manipulation DetectionJinjie Shen, Yaxiong Wang, Yujiao Wu, Lechao Cheng 等ICML 2026
