IDseq: Decoupled and Sequentially Detecting and Grounding Multi-Modal Media Manipulation
Runxin Liu, Tian Xie, Jiaming Li, Lingyun Yu, Hongtao Xie
摘要
Detecting and grounding multi-modal media manipulation aims to categorize the type and localize the region of manipulation for image-text pairs in both two modalities. Existing methods have not sufficiently explored the intrinsic properties of the manipulated images, which contain both forgery and content features, leading to inefficient utilization. To address this problem, we propose an Image-Driven Decoupled Sequential Framework (IDseq), designed to decouple image features and rationally integrate them to accomplish different sub-tasks effectively. Specifically, IDseq employs two specially designed disentangled losses to guide the disentangled learning of forgery and content features. To efficiently leverage these features, we propose a Decoupled Image Manipulation Decoder (DIMD) that processes image tasks within a decoupled schema. We mitigate their exclusive competition by separating the image tasks into forgery-relevant and content-relevant components and training them without gradient interaction. Additionally, we utilize content features enhanced by the proposed Manipulation Indicator Generator (MIG) for the text tasks, which provide the maximal visual information as a reference while eliminating interference from unverified image data. Extensive experiments show the superiority of our IDseq, where it notably outperforms SOTA methods on the fine-grained classification by 3.8% in mAP and the forgery face grounding by 8.7% in IoUmean, even 1.3% in F1 on the most challenging manipulated text grounding.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- StyleCLIP: Text-Driven Manipulation of StyleGAN ImageryOr Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or 等ICCV 2021 · 被引用 1,437 次
相关 Paper
- Beyond Artificial Misalignment: Detecting and Grounding Semantic-Coordinated Multimodal ManipulationsJinjie Shen, Yaxiong Wang, Lechao Cheng, Nan Pu 等ACM MM 2025 · 被引用 2 次
- Critical Forgetting-Based Multi-Scale Disentanglement for Deepfake DetectionKai Li, Wenqi Ren, Jianshu Li, Wei Wang 等AAAI 2025 · 被引用 3 次
- ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and GroundingZhenxing Zhang, Yaxiong Wang, Lechao Cheng, Zhun Zhong 等CVPR 2025
- UCF: Uncovering Common Features for Generalizable Deepfake DetectionZhiyuan Yan, Yong Zhang, Yanbo Fan, Baoyuan WuICCV 2023 · 被引用 264 次
- Unleashing the Potential of Consistency Learning for Detecting and Grounding Multi-Modal Media ManipulationYiheng Li, Yang Yang, Zichang Tan, Huan Liu 等CVPR 2025
