A Multimodal Deviation Perceiving Framework for Weakly-Supervised Temporal Forgery Localization
Wenbo Xu, Junyan Wu, Wei Lu, Xiangyang Luo, Qian Wang
Abstract
Current researches on Deepfake forensics often treat detection as a classification task or temporal forgery localization problem, which are usually restrictive, time-consuming, and challenging to scale for large datasets. To resolve these issues, we present a multimodal deviation perceiving framework for weakly-supervised temporal forgery localization (MDP), which aims to identify temporal partial forged segments using only video-level annotations. The MDP proposes a novel multimodal interaction mechanism (MI) and an extensible deviation perceiving loss to perceive multimodal deviation, which achieves the refined start and end timestamps localization of forged segments. Specifically, MI introduces a temporal property preserving cross-modal attention to measure the relevance between the visual and audio modalities in the probabilistic embedding space. It could identify the inter-modality deviation and construct comprehensive video features for temporal forgery localization. To explore further temporal deviation for weakly-supervised learning, an extensible deviation perceiving loss has been proposed, aiming at enlarging the deviation of adjacent segments of the forged samples and reducing that of genuine samples. Extensive experiments demonstrate the effectiveness of the proposed framework and achieve comparable results to fully-supervised approaches in several evaluation metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a31bdf51-c417-479e-af4a-25025a24543eCited by top-tier papers1
Ask how each one uses itBuilds on14
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Joint Audio-Visual Deepfake DetectionYipin Zhou, Ser-Nam LimICCV 2021 · 232 citations
- Not made for each other- Audio-Visual Dissonance-based Deepfake Detection and LocalizationKomal Chugh, Parul Gupta, Abhinav Dhall, Ramanathan SubramanianACM MM 2020 · 217 citations
- AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake DatasetZhixi Cai, Shreya Ghosh, Aman Pankaj Adatia, Munawar Hayat et al.ACM MM 2024 · 51 citations
Related papers
- Structural–Semantic Perception for Diffusion-Guided Temporal Forgery LocalizationLigong Cao, Yeting Guo, Haoang ChiCVPR 2026
- Intra-Modal and Cross-Modal Synchronization for Audio-Visual Deepfake Detection and Temporal LocalizationAshutosh Anshul, Shreyas Gopal, Deepu Rajan, Eng Siong ChngICCV 2025 · 10 citations
- FRADE: Forgery-aware Audio-distilled Multimodal Learning for Deepfake DetectionFan Nie, Jiangqun Ni, Jian Zhang, Bin Zhang et al.ACM MM 2024 · 17 citations
- Coarse-to-Fine Proposal Refinement Framework for Audio Temporal Forgery Detection and LocalizationJunyan Wu, Wei Lu, Xiangyang Luo, Rui Yang et al.ACM MM 2024 · 16 citations
- Query-Based Audio-Visual Temporal Forgery Localization with Register-Enhanced Representation LearningXiaodong Zhu, Suting Wang, Junqi Yang, Yuhong Yang et al.ACM MM 2025
