RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization
Wen Huang, Jiarui Yang, Tao Dai, Jiawei Li, Shaoxiong Zhan, Bin Wang, Shu-Tao Xia
Abstract
Visual manipulation localization (VML) aims to identify tampered regions in images and videos, a task that has become increasingly challenging with the rise of advanced editing tools. Existing methods face two central issues. The first is resolution diversity. Resizing or padding can distort subtle forensic cues and introduce unnecessary computational cost. The second is the difficulty of extending spatial models for images to spatio-temporal inputs in videos, which often results in maintaining separate architectures for the two data types. To address these challenges, we propose RelayFormer, a unified framework that adapts to varying resolutions and naturally handles both static and temporal visual data. RelayFormer partitions inputs into fixed-size sub-images and introduces Global Local Relay (GLR) tokens that propagate structured context through a relay-based attention mechanism. This design enables efficient exchange of global cues, such as semantic or temporal consistency, while preserving fine-grained manipulation artifacts. Unlike prior approaches that depend on uniform resizing or sparse attention, RelayFormer scales to variable resolutions and video sequences with minimal overhead. Experiments across diverse benchmarks demonstrate superior performance and strong efficiency, combining resolution adaptivity without interpolation or excessive padding, unified processing for images and videos, and a favorable balance between accuracy and computational cost. Code is available at https://github.com/WenOOI/RelayFormer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ba54607-e44f-497c-ac9f-2f68e544162eBuilds on11
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Image Manipulation Detection by Multi-View Multi-Scale SupervisionXinru Chen, Chengbo Dong, Jiaqi Ji, Juan Cao et al.ICCV 2021 · 271 citations
- MOSE: A New Dataset for Video Object Segmentation in Complex ScenesHenghui Ding, Chang Liu, Shuting He, Xudong Jiang et al.ICCV 2023 · 267 citations
- ObjectFormer for Image Manipulation Detection and LocalizationJunke Wang, Zuxuan Wu, Jingjing Chen, Xintong Han et al.CVPR 2022 · 190 citations
- FuseFormer: Fusing Fine-Grained Information in Transformers for Video InpaintingRui Liu, Hanming Deng, Yangyi Huang, Xiaoyu Shi et al.ICCV 2021 · 165 citations
Related papers
- M2sformer: Multi-Spectral and Multi-Scale Attention With Edge-Aware Difficulty Guidance for Image Forgery LocalizationJu-Hyeon Nam, Dong-Hyun Moon, Sang-Chul LeeICCV 2025 · 4 citations
- TransForensics: Image Forgery Localization with Dense Self-AttentionJing Hao, Zhixin Zhang, Shicai Yang, Di Xie et al.ICCV 2021 · 77 citations
- DeformTrace: A Deformable State Space Model with Relay Tokens for Temporal Forgery LocalizationXiaodong Zhu, Suting Wang, Yuanming Zheng, Junqi Yang et al.AAAI 2026
- Collaborative Transformers with Multi-Level Forensic Attention for Image Manipulation LocalizationJiwei Zhang, Wenbo Feng, Siwei Wang, Feifei Kou et al.AAAI 2026
- Former: Unified Retrieval and Reranking Transformer for Place RecognitionSijie Zhu, Linjie Yang, Chen Chen, Mubarak Shah et al.CVPR 2023
