Frequency Mining Empowered by Text Aggregation: A New Perspective on Document Image Tampering Detection
Ziqi Yi, Guitao Xu, Shihang Wu, Peirong Zhang, Lianwen Jin
Abstract
Document image tampering detection faces significant challenges due to the subtle and spatially dispersed nature of tampering traces, which are often confined to localized regions within tampered text. While existing methods leverage frequency domain information to reveal hidden artifacts, they fail to fully exploit the rich frequency spectrum and lack effective mechanisms for aggregating scattered tampering evidence across extended text regions. To overcome these limitations, we propose the Text Aggregation and multi-Frequency Enhancement Network (TAFE-Net). Specifically, to capture more subtle tampering traces, we design a Multi-Frequency Feature Extractor that comprehensively utilizes various proven effective frequency information. In addition, the Visual-Frequency Integration Module and Direction-aware Frequency Decoupling Enhancement module are introduced to aggregate text features in both horizontal and vertical directions within the frequency domain, from coarse to fine granularity, addressing the incomplete detection of tampered text caused by dispersed tampering traces. Experiments on the DocTamper and RTM datasets demonstrate that our approach establishes new state-of-the-art results and maintains superior robustness against various degradations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on12
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- ObjectFormer for Image Manipulation Detection and LocalizationJunke Wang, Zuxuan Wu, Jingjing Chen, Xintong Han et al.CVPR 2022 · 190 citations
- Mesoscopic Insights: Orchestrating Multi-Scale & Hybrid Architecture for Image Manipulation LocalizationXuekang Zhu, Xiaochen Ma, Lei Su, Zhuohang Jiang et al.AAAI 2025 · 44 citations
- Can We Get Rid of Handcrafted Feature Extractors? SparseViT: Nonsemantics-Centered, Parameter-Efficient Image Manipulation Localization Through Spare-Coding TransformerLei Su, Xiaochen Ma, Xuekang Zhu, Chaoqun Niu et al.AAAI 2025 · 10 citations
Related papers
- Towards Robust Tampered Text Detection in Document Image: New Dataset and New SolutionChenfan Qu, Chongyu Liu, Yuliang Liu, Xinhong Chen et al.CVPR 2023
- ADCD-Net: Robust Document Image Forgery Localization via Adaptive DCT Feature and Hierarchical Content DisentanglementKahim Wong, Jicheng Zhou, Haiwei Wu, Yain-Whar Si et al.ICCV 2025 · 3 citations
- D2FANet: Enhancing Video Object Detection with Dual-Domain Feature Aggregation NetworkQiang Qi, Wenqi Shang, Meifang Wang, Xiao WangCVPR 2026
- Exploiting Fine-Grained Face Forgery Clues via Progressive Enhancement LearningQiqi Gu, Shen Chen, Taiping Yao, Yang Chen et al.AAAI 2022 · 187 citations
- Multi-Attentional Deepfake DetectionHanqing Zhao, Wenbo Zhou, Dongdong Chen, Tianyi Wei et al.CVPR 2021
