TextShield-R1: Reinforced Reasoning for Tampered Text Detection
Chenfan Qu, Yiwu Zhong, Jian Liu, Xuekang Zhu, Bohan Yu, Lianwen Jin
Abstract
The growing prevalence of tampered images poses serious security threats, highlighting the urgent need for reliable detection methods. Multimodal large language models (MLLMs) demonstrate strong potential in analyzing tampered images and generating interpretations. However, they still struggle with identifying micro-level artifacts, exhibit low accuracy in localizing tampered text regions, and heavily rely on expensive annotations for forgery interpretation. To this end, we introduce TextShield-R1, the first reinforcement learning based MLLM solution for tampered text detection and reasoning. Specifically, our approach introduces Forensic Continual Pre-training, an easy-to-hard curriculum that well prepares the MLLM for tampered text detection by harnessing the large-scale cheap data from natural image forensic and OCR tasks. During fine-tuning, we perform Group Relative Policy Optimization with novel reward functions to reduce annotation dependency and improve reasoning capabilities. At inference time, we enhance localization accuracy via OCR Rectification, a method that leverages the MLLM's strong text recognition abilities to refine its predictions. Furthermore, to support rigorous evaluation, we introduce the Text Forensics Reasoning (TFR) benchmark, comprising over 45k real and tampered images across 16 languages, 10 tampering techniques, and diverse domains. Rich reasoning-style annotations are included, allowing for comprehensive assessment. Our TFR benchmark simultaneously addresses seven major limitations of existing benchmarks and enables robust evaluation under cross-style, cross-method, and cross-language conditions. Extensive experiments demonstrate that TextShield-R1 significantly advances the state of the art in interpretable tampered text detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ebaa39e0-7f8b-4c8d-9f6c-88fa778eb7e0Cited by top-tier papers1
- Detect Any AI-Counterfeited Text ImageChenfan Qu, Yiwu Zhong, Xuekang Zhu, Junchi Li et al.CVPR 2026
Builds on16
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- DiffUTE: Universal Text Editing Diffusion ModelHaoxing Chen, Zhuoer Xu, Zhangxuan Gu, Jun Lan et al.NeurIPS 2023 · 61 citations
- Mesoscopic Insights: Orchestrating Multi-Scale & Hybrid Architecture for Image Manipulation LocalizationXuekang Zhu, Xiaochen Ma, Lei Su, Zhuohang Jiang et al.AAAI 2025 · 44 citations
- Revisiting Tampered Scene Text Detection in the Era of Generative AIChenfan Qu, Yiwu Zhong, Fengjun Guo, Lianwen JinAAAI 2025 · 20 citations
Related papers
- FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded ReasoningYikun Ji, Yan Hong, Qi Fan, Jun Lan et al.ICLR 2026 · 9 citations
- From Pixels to Semantics: A Novel MLLM-Driven Approach for Explainable Tampered Text DetectionGuitao Xu, Ziqi Yi, Peirong Zhang, Jiahuan Cao et al.ACM MM 2025 · 2 citations
- ForgerySleuth: Empowering Multimodal Large Language Models for Image Manipulation DetectionZhihao Sun, Haoran Jiang, Haoran Chen, Yixin Cao et al.NeurIPS 2025 · 16 citations
- FakeShield: Explainable Image Forgery Detection and Localization via Multi-modal Large Language ModelsZhipei Xu, Xuanyu Zhang, Runyi Li, Zecheng Tang et al.ICLR 2025
- Reasoning-Driven Anomaly Detection and Localization with Image-Level SupervisionYizhou Jin, Yuezhu Feng, Jinjin Zhang, Peng Wang et al.CVPR 2026 · 4 citations
