MGQFormer: Mask-Guided Query-Based Transformer for Image Manipulation Localization
Kunlun Zeng, Ri Cheng, Weimin Tan, Bo Yan
Abstract
Deep learning-based models have made great progress in image tampering localization, which aims to distinguish between manipulated and authentic regions. However, these models suffer from inefficient training. This is because they use ground-truth mask labels mainly through the cross-entropy loss, which prioritizes per-pixel precision but disregards the spatial location and shape details of manipulated regions. To address this problem, we propose a Mask-Guided Query-based Transformer Framework (MGQFormer), which uses ground-truth masks to guide the learnable query token (LQT) in identifying the forged regions. Specifically, we extract feature embeddings of ground-truth masks as the guiding query token (GQT) and feed GQT and LQT into MGQFormer to estimate fake regions, respectively. Then we make MGQFormer learn the position and shape information in ground-truth mask labels by proposing a mask-guided loss to reduce the feature distance between GQT and LQT. We also observe that such mask-guided training strategy has a significant impact on the convergence speed of MGQFormer training. Extensive experiments on multiple benchmarks show that our method significantly improves over state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2f868069-8cc5-4c6b-ace8-face6e6631f3Cited by top-tier papers5
- Mesoscopic Insights: Orchestrating Multi-Scale & Hybrid Architecture for Image Manipulation LocalizationXuekang Zhu, Xiaochen Ma, Lei Su, Zhuohang Jiang et al.AAAI 2025 · 44 citations
- SAFIRE: Segment Any Forged Image RegionMyung-Joon Kwon, Wonjun Lee, Seung-Hun Nam, Minji Son et al.AAAI 2025 · 25 citations
- MUN: Image Forgery Localization Based on M³ Encoder and UN DecoderYaqi Liu, Shuhuan Chen, Haichao Shi, Xiaoyu Zhang et al.AAAI 2025 · 6 citations
- Weakly-Supervised Image Forgery Localization via Vision-Language Collaborative Reasoning FrameworkZiqi Sheng, Junyan Wu, Wei Lu, Jiantao ZhouAAAI 2026 · 3 citations
- M²RL-Net: Multi-View and Multi-Level Relation Learning Network for Weakly-Supervised Image Forgery DetectionJiafeng Li, Ying Wen, Lianghua HeAAAI 2025 · 2 citations
Builds on11
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Segmenter: Transformer for Semantic SegmentationRobin Strudel, Ricardo Garcia, Ivan Laptev, Cordelia SchmidICCV 2021 · 1,898 citations
- Image Manipulation Detection by Multi-View Multi-Scale SupervisionXinru Chen, Chengbo Dong, Jiaqi Ji, Juan Cao et al.ICCV 2021 · 271 citations
Related papers
- M2sformer: Multi-Spectral and Multi-Scale Attention With Edge-Aware Difficulty Guidance for Image Forgery LocalizationJu-Hyeon Nam, Dong-Hyun Moon, Sang-Chul LeeICCV 2025 · 4 citations
- A Unified Query-based Paradigm for Camouflaged Instance SegmentationBo Dong, Jialun Pei, Rongrong Gao, Tian-Zhu Xiang et al.ACM MM 2023 · 19 citations
- TransForensics: Image Forgery Localization with Dense Self-AttentionJing Hao, Zhixin Zhang, Shicai Yang, Di Xie et al.ICCV 2021 · 77 citations
- MP-Former: Mask-Piloted Transformer for Image SegmentationHao Zhang, Feng Li, Huaizhe Xu, Shijia Huang et al.CVPR 2023
- Beyond [CLS] Token: Query-Driven Token-Level Forgery Purification for Generalizable Deepfake DetectionChangshuo Wang, Jiangming Wang, Ke-Yue Zhang, Taiping Yao et al.CVPR 2026
