Intra-Modal and Cross-Modal Synchronization for Audio-Visual Deepfake Detection and Temporal Localization
Ashutosh Anshul, Shreyas Gopal, Deepu Rajan, Eng Siong Chng
Abstract
Recent deepfake detection algorithms focus solely on uni-modal or cross-modal inconsistencies. While the former disregards audio-visual correspondence entirely rendering them less effective against multimodal attacks, the latter overlooks inconsistencies in a particular modality. Moreover, many models are single-stage supervised frameworks, effective on specific training data but less generalizable to new manipulations. To address these gaps, we propose a two-stage multimodal framework that first learns intra-modal and cross-modal temporal synchronization on real videos, capturing audio-visual correspondences crucial for deepfake detection and localization. We introduce a Gaussian-targeted loss in our pretraining model to focus on learning relative synchronization patterns across multimodal pairs. Using pretrained features, our approach not only enables classification on fully manipulated videos but also supports a localization module for partial deepfakes with only specific segments spoofed. Moreover, the pretraining stage does not require fine-tuning, thus reducing complexity. Our model, tested on various benchmark datasets, demonstrates strong generalization and precise temporal localization. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b6455df-d42e-42bd-8f06-bc93292c49ddCited by top-tier papers1
Ask how each one uses itBuilds on29
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess et al.ICCV 2019 · 2,966 citations
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- FSGAN: Subject Agnostic Face Swapping and ReenactmentYuval Nirkin, Yosi Keller, Tal HassnerICCV 2019 · 710 citations
- BMN: Boundary-Matching Network for Temporal Action Proposal GenerationTianwei Lin, Xiao Liu, Xin Li, Errui Ding et al.ICCV 2019 · 709 citations
Related papers
- Multi-modal Deepfake Detection via Multi-task Audio-Visual Prompt LearningHui Miao, Yuanfang Guo, Zeming Liu, Yunhong WangAAAI 2025 · 8 citations
- AVFF: Audio-Visual Feature Fusion for Video Deepfake DetectionTrevine Oorloff, Surya Koppisetti, Nicolò Bonettini, Divyaraj Solanki et al.CVPR 2024 · 51 citations
- A Multimodal Deviation Perceiving Framework for Weakly-Supervised Temporal Forgery LocalizationWenbo Xu, Junyan Wu, Wei Lu, Xiangyang Luo et al.ACM MM 2025 · 2 citations
- Not made for each other- Audio-Visual Dissonance-based Deepfake Detection and LocalizationKomal Chugh, Parul Gupta, Abhinav Dhall, Ramanathan SubramanianACM MM 2020 · 217 citations
- Audio-Visual Asynchrony Mitigation: Cross-Modal Alignment and Feature Reconstruction for Deepfake DetectionYan Wang, Qindong Sun, Dongzhu RongACM MM 2025 · 1 citation
