MGQFormer: Mask-Guided Query-Based Transformer for Image Manipulation Localization
Kunlun Zeng, Ri Cheng, Weimin Tan, Bo Yan
摘要
Deep learning-based models have made great progress in image tampering localization, which aims to distinguish between manipulated and authentic regions. However, these models suffer from inefficient training. This is because they use ground-truth mask labels mainly through the cross-entropy loss, which prioritizes per-pixel precision but disregards the spatial location and shape details of manipulated regions. To address this problem, we propose a Mask-Guided Query-based Transformer Framework (MGQFormer), which uses ground-truth masks to guide the learnable query token (LQT) in identifying the forged regions. Specifically, we extract feature embeddings of ground-truth masks as the guiding query token (GQT) and feed GQT and LQT into MGQFormer to estimate fake regions, respectively. Then we make MGQFormer learn the position and shape information in ground-truth mask labels by proposing a mask-guided loss to reduce the feature distance between GQT and LQT. We also observe that such mask-guided training strategy has a significant impact on the convergence speed of MGQFormer training. Extensive experiments on multiple benchmarks show that our method significantly improves over state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Mesoscopic Insights: Orchestrating Multi-Scale & Hybrid Architecture for Image Manipulation LocalizationXuekang Zhu, Xiaochen Ma, Lei Su, Zhuohang Jiang 等AAAI 2025 · 被引用 44 次
- SAFIRE: Segment Any Forged Image RegionMyung-Joon Kwon, Wonjun Lee, Seung-Hun Nam, Minji Son 等AAAI 2025 · 被引用 25 次
- MUN: Image Forgery Localization Based on M³ Encoder and UN DecoderYaqi Liu, Shuhuan Chen, Haichao Shi, Xiaoyu Zhang 等AAAI 2025 · 被引用 6 次
- Weakly-Supervised Image Forgery Localization via Vision-Language Collaborative Reasoning FrameworkZiqi Sheng, Junyan Wu, Wei Lu, Jiantao ZhouAAAI 2026 · 被引用 3 次
- M²RL-Net: Multi-View and Multi-Level Relation Learning Network for Weakly-Supervised Image Forgery DetectionJiafeng Li, Ying Wen, Lianghua HeAAAI 2025 · 被引用 2 次
它引用的顶会 Paper11
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- Segmenter: Transformer for Semantic SegmentationRobin Strudel, Ricardo Garcia, Ivan Laptev, Cordelia SchmidICCV 2021 · 被引用 1,898 次
- Image Manipulation Detection by Multi-View Multi-Scale SupervisionXinru Chen, Chengbo Dong, Jiaqi Ji, Juan Cao 等ICCV 2021 · 被引用 271 次
相关 Paper
- M2sformer: Multi-Spectral and Multi-Scale Attention With Edge-Aware Difficulty Guidance for Image Forgery LocalizationJu-Hyeon Nam, Dong-Hyun Moon, Sang-Chul LeeICCV 2025 · 被引用 4 次
- A Unified Query-based Paradigm for Camouflaged Instance SegmentationBo Dong, Jialun Pei, Rongrong Gao, Tian-Zhu Xiang 等ACM MM 2023 · 被引用 19 次
- TransForensics: Image Forgery Localization with Dense Self-AttentionJing Hao, Zhixin Zhang, Shicai Yang, Di Xie 等ICCV 2021 · 被引用 77 次
- MP-Former: Mask-Piloted Transformer for Image SegmentationHao Zhang, Feng Li, Huaizhe Xu, Shijia Huang 等CVPR 2023
- Beyond [CLS] Token: Query-Driven Token-Level Forgery Purification for Generalizable Deepfake DetectionChangshuo Wang, Jiangming Wang, Ke-Yue Zhang, Taiping Yao 等CVPR 2026
