Collaborative Transformers with Multi-Level Forensic Attention for Image Manipulation Localization
Jiwei Zhang, Wenbo Feng, Siwei Wang, Feifei Kou, Haoyang Yu, Shaozhang Niu
Abstract
The proliferation of the tampered images on social media can pose serious societal risks, influencing public opinion and causing panic. Image Manipulation Localization technique has advanced to address this, but some methods focus on microscopic traces, overlooking macroscopic semantics that deceive viewers. To address this problem, we propose a novel Image Manipulation Localization framework called Collaborative Transformers (Co-Transformers), designed to fully explore and utilize the collaborative information between macroscopic semantics and microscopic traces. This framework is based on two Vision Transformer variants. The first variant captures the semantic logic of the image. The second variant delves into microscopic tampering traces. By dynamically fusing these two complementary features, the framework enables interaction between macroscopic semantic inconsistencies and microscopic abnormal traces, effectively coordinating their relationship in the latent space. Furthermore, we introduce a new Multi-Level Forensic Attention (MLF-Attention) mechanism to enhance the model's ability to extract various tampered traces, this mechanism can be integrated into our framework. Compared with existing methods, our proposed framework achieves state-of-the-art results in localization accuracy and shows good robustness against various attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 52be7cf2-04fe-4797-be09-1567af682270Builds on7
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- ObjectFormer for Image Manipulation Detection and LocalizationJunke Wang, Zuxuan Wu, Jingjing Chen, Xintong Han et al.CVPR 2022 · 190 citations
- Mesoscopic Insights: Orchestrating Multi-Scale & Hybrid Architecture for Image Manipulation LocalizationXuekang Zhu, Xiaochen Ma, Lei Su, Zhuohang Jiang et al.AAAI 2025 · 44 citations
- JPEG Compression-aware Image Forgery LocalizationMenglu Wang, Xueyang Fu, Jiawei Liu, Zheng-Jun ZhaACM MM 2022 · 22 citations
Related papers
- TransForensics: Image Forgery Localization with Dense Self-AttentionJing Hao, Zhixin Zhang, Shicai Yang, Di Xie et al.ICCV 2021 · 77 citations
- Edge-aware Affinity Enhancement for Image Manipulation LocalizationTianyi Zhang, Qinglong Lin, Yang Hu, Pengming Feng et al.ACM MM 2025 · 1 citation
- Generic Image Manipulation Localization through the Lens of Multi-scale Spatial InconsistenceZan Gao, Shenghao Chen, Yangyang Guo, Weili Guan et al.ACM MM 2022 · 14 citations
- M2sformer: Multi-Spectral and Multi-Scale Attention With Edge-Aware Difficulty Guidance for Image Forgery LocalizationJu-Hyeon Nam, Dong-Hyun Moon, Sang-Chul LeeICCV 2025 · 4 citations
- EARG-Net: Edge-Aware Reconstruction-Guided Network for Image Manipulation Detection and LocalizationYanpu Yu, Zhaoxin Shi, Hanqing Zhao, Tianyi Wei et al.AAAI 2026 · 1 citation
