Not made for each other- Audio-Visual Dissonance-based Deepfake Detection and Localization
Komal Chugh, Parul Gupta, Abhinav Dhall, Ramanathan Subramanian
Abstract
We propose detection of deepfake videos based on the dissimilarity between the audio and visual modalities, termed as the Modality Dissonance Score (MDS). We hypothesize that manipulation of either modality will lead to dis-harmony between the two modalities, e.g., loss of lip-sync, unnatural facial and lip movements, etc. MDS is computed as the mean aggregate of dissimilarity scores between audio and visual segments in a video. Discriminative features are learnt for the audio and visual channels in a chunk-wise manner, employing the cross-entropy loss for individual modalities, and a contrastive loss that models inter-modality similarity. Extensive experiments on the DFDC and DeepFake-TIMIT Datasets show that our approach outperforms the state-of-the-art by up to 7%. We also demonstrate temporal forgery localization, and show how our technique identifies the manipulated video segments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e62ef4e0-7c80-482c-a395-183f8c54a134Cited by top-tier papers19
- Leveraging Real Talking Faces via Self-Supervision for Robust Forgery DetectionAlexandros Haliassos, Rodrigo Mira, Stavros Petridis, Maja PanticCVPR 2022 · 138 citations
- Exposing the Deception: Uncovering More Forgery Clues for Deepfake DetectionZhongjie Ba, Qingyu Liu, Zhenguang Liu, Shuang Wu et al.AAAI 2024 · 101 citations
- Learning Second Order Local Anomaly for General Face Forgery DetectionJianwei Fei, Yunshu Dai, Peipeng Yu, Tianrun Shen et al.CVPR 2022 · 75 citations
- Lips Are Lying: Spotting the Temporal Inconsistency between Audio and Visual in Lip-Syncing DeepFakesWeifeng Liu, Tianyi She, Jiawei Liu, Boheng Li et al.NeurIPS 2024 · 57 citations
- AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake DatasetZhixi Cai, Shreya Ghosh, Aman Pankaj Adatia, Munawar Hayat et al.ACM MM 2024 · 51 citations
Builds on2
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess et al.ICCV 2019 · 2,966 citations
- Emotions Don't Lie: An Audio-Visual Deepfake Detection Method using Affective CuesTrisha Mittal, Uttaran Bhattacharya, Rohan Chandra, Aniket Bera et al.ACM MM 2020 · 314 citations
Related papers
- AVFF: Audio-Visual Feature Fusion for Video Deepfake DetectionTrevine Oorloff, Surya Koppisetti, Nicolò Bonettini, Divyaraj Solanki et al.CVPR 2024 · 51 citations
- Joint Audio-Visual Deepfake DetectionYipin Zhou, Ser-Nam LimICCV 2021 · 232 citations
- A Multimodal Deviation Perceiving Framework for Weakly-Supervised Temporal Forgery LocalizationWenbo Xu, Junyan Wu, Wei Lu, Xiangyang Luo et al.ACM MM 2025 · 2 citations
- Audio-Visual Asynchrony Mitigation: Cross-Modal Alignment and Feature Reconstruction for Deepfake DetectionYan Wang, Qindong Sun, Dongzhu RongACM MM 2025 · 1 citation
- Intra-Modal and Cross-Modal Synchronization for Audio-Visual Deepfake Detection and Temporal LocalizationAshutosh Anshul, Shreyas Gopal, Deepu Rajan, Eng Siong ChngICCV 2025 · 10 citations
