UMMAFormer: A Universal Multimodal-adaptive Transformer Framework for Temporal Forgery Localization
Rui Zhang, Hongxia Wang, Mingshan Du, Hanqing Liu, Yang Zhou, Qiang Zeng
Abstract
The emergence of artificial intelligence-generated content (AIGC) has raised concerns about the authenticity of multimedia content in various fields. However, existing research for forgery content detection has focused mainly on binary classification tasks of complete videos, which has limited applicability in industrial settings. To address this gap, we propose UMMAFormer, a novel universal transformer framework for temporal forgery localization (TFL) that predicts forgery segments with multimodal adaptation. Our approach introduces a Temporal Feature Abnormal Attention (TFAA) module based on temporal feature reconstruction to enhance the detection of temporal differences. We also design a Parallel Cross-Attention Feature Pyramid Network (PCA-FPN) to optimize the Feature Pyramid Network (FPN) for subtle feature enhancement. To evaluate the proposed method, we contribute a novel Temporal Video Inpainting Localization (TVIL) dataset specifically tailored for video inpainting scenes. Our experiments show that our approach achieves state-of-the-art performance on benchmark datasets, including Lav-DF, TVIL, and Psynd, significantly outperforming previous methods. The code and data are available at https://github.com/ymhzyj/UMMAFormer/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 892e6226-7d99-4193-b8e0-c14c73ce93f1Cited by top-tier papers8
- AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake DatasetZhixi Cai, Shreya Ghosh, Aman Pankaj Adatia, Munawar Hayat et al.ACM MM 2024 · 51 citations
- Coarse-to-Fine Proposal Refinement Framework for Audio Temporal Forgery Detection and LocalizationJunyan Wu, Wei Lu, Xiangyang Luo, Rui Yang et al.ACM MM 2024 · 16 citations
- Intra-Modal and Cross-Modal Synchronization for Audio-Visual Deepfake Detection and Temporal LocalizationAshutosh Anshul, Shreyas Gopal, Deepu Rajan, Eng Siong ChngICCV 2025 · 10 citations
- ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in VideosPeijun Bao, Anwei Luo, Gang Pan, Alex C. Kot et al.CVPR 2026 · 2 citations
- A Multimodal Deviation Perceiving Framework for Weakly-Supervised Temporal Forgery LocalizationWenbo Xu, Junyan Wu, Wei Lu, Xiangyang Luo et al.ACM MM 2025 · 2 citations
Builds on24
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess et al.ICCV 2019 · 2,966 citations
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
- WildDeepfake: A Challenging Real-World Dataset for Deepfake DetectionBojia Zi, Minghao Chang, Jingjing Chen, Xingjun Ma et al.ACM MM 2020 · 443 citations
- UniT: Multimodal Multitask Learning with a Unified TransformerRonghang Hu, Amanpreet SinghICCV 2021 · 354 citations
Related papers
- On Learning Multi-Modal Forgery Representation for Diffusion Generated Video DetectionXiufeng Song, Xiao Guo, Jiache Zhang, Qirui Li et al.NeurIPS 2024 · 63 citations
- Query-Based Audio-Visual Temporal Forgery Localization with Register-Enhanced Representation LearningXiaodong Zhu, Suting Wang, Junqi Yang, Yuhong Yang et al.ACM MM 2025
- Beyond Spatial Frequency: Pixel-Wise Temporal Frequency-Based Deepfake Video DetectionTaehoon Kim, Jongwook Choi, Yonghyun Jeong, Haeun Noh et al.ICCV 2025 · 7 citations
- FRADE: Forgery-aware Audio-distilled Multimodal Learning for Deepfake DetectionFan Nie, Jiangqun Ni, Jian Zhang, Bin Zhang et al.ACM MM 2024 · 17 citations
- Unlocking the Capabilities of Large Vision-Language Models for Generalizable and Explainable Deepfake DetectionPeipeng Yu, Jianwei Fei, Hui Gao, Xuan Feng et al.ICML 2025
