Multi-scale Target-Aware Framework for Constrained Splicing Detection and Localization
Yuxuan Tan, Yuanman Li, Limin Zeng, Jiaxiong Ye, Wei Wang, Xia Li
Abstract
Constrained image splicing detection and localization (CISDL) is a fundamental task of multimedia forensics, which detects splicing operation between two suspected images and localizes the spliced region on both images. Recent works regard it as a deep matching problem and have made significant progress. However, existing frameworks typically perform feature extraction and correlation matching as separate processes, which may hinder the model's ability to learn discriminative features for matching and can be susceptible to interference from ambiguous background pixels. In this work, we propose a multi-scale target-aware framework to couple feature extraction and correlation matching in a unified pipeline. In contrast to previous methods, we design a target-aware attention mechanism that jointly learns features and performs correlation matching between the probe and donor images. Our approach can effectively promote the collaborative learning of related patches, and perform mutual promotion of feature learning and correlation matching. Additionally, in order to handle scale transformations, we introduce a multi-scale projection method, which can be readily integrated into our target-aware framework that enables the attention process to be conducted between tokens containing information of varying scales. Our experiments demonstrate that our model, which uses a unified pipeline, outperforms state-of-the-art methods on several benchmark datasets and is robust against scale transformations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on5
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 746 citations
Related papers
- Weakly-Supervised Video Re-Localization with Multiscale Attention ModelYung-Han Huang, Kuang-Jui Hsu, Shyh-Kang Jeng, Yen-Yu LinAAAI 2020 · 12 citations
- On the Detection of Digital Face ManipulationHao Dang, Feng Liu, Joel Stehouwer, Xiaoming Liu et al.CVPR 2020
- M²RL-Net: Multi-View and Multi-Level Relation Learning Network for Weakly-Supervised Image Forgery DetectionJiafeng Li, Ying Wen, Lianghua HeAAAI 2025 · 2 citations
- MatchDet: A Collaborative Framework for Image Matching and Object DetectionJinxiang Lai, Wenlong Wu, Bin-Bin Gao, Jun Liu et al.AAAI 2024 · 1 citation
- Contextually-Guided State Space Fusion for Misaligned Multi-Spectral Object DetectionGuyue Jin, Tianming Zhao, Jiacan Yan, Tian TianACM MM 2025 · 1 citation
