Emotions Don't Lie: An Audio-Visual Deepfake Detection Method using Affective Cues
Trisha Mittal, Uttaran Bhattacharya, Rohan Chandra, Aniket Bera, Dinesh Manocha
摘要
We present a learning-based method for detecting real and fake deepfake multimedia content. To maximize information for learning, we extract and analyze the similarity between the two audio and visual modalities from within the same video. Additionally, we extract and compare affective cues corresponding to perceived emotion from the two modalities within a video to infer whether the input video is "real" or "fake". We propose a deep learning network, inspired by the Siamese network architecture and the triplet loss. To validate our model, we report the AUC metric on two large-scale deepfake detection datasets, DeepFake-TIMIT Dataset and DFDC. We compare our approach with several SOTA deepfake detection methods and report per-video AUC of 84.4% on the DFDC and 96.6% on the DF-TIMIT datasets, respectively. To the best of our knowledge, ours is the first approach that simultaneously exploits audio and video modalities and also perceived emotions from the two modalities for deepfake detection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Not made for each other- Audio-Visual Dissonance-based Deepfake Detection and LocalizationKomal Chugh, Parul Gupta, Abhinav Dhall, Ramanathan SubramanianACM MM 2020 · 被引用 217 次
- AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake DatasetZhixi Cai, Shreya Ghosh, Aman Pankaj Adatia, Munawar Hayat 等ACM MM 2024 · 被引用 51 次
- Contrastive Learning of Global and Local Video RepresentationsShuang Ma, Zhaoyang Zeng, Daniel McDuff, Yale SongNeurIPS 2021 · 被引用 45 次
- Combating Online Misinformation Videos: Characterization, Detection, and Future DirectionsYuyan Bu, Qiang Sheng, Juan Cao, Peng Qi 等ACM MM 2023 · 被引用 38 次
- "It Matches My Worldview": Examining Perceptions and Attitudes Around Fake VideosFarhana Shahid, Srujana Kamath, Annie Sidotam, Vivian Jiang 等CHI 2022 · 被引用 38 次
它引用的顶会 Paper4
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess 等ICCV 2019 · 被引用 2,966 次
- M3ER: Multiplicative Multimodal Emotion Recognition using Facial, Textual, and Speech CuesTrisha Mittal, Uttaran Bhattacharya, Rohan Chandra, Aniket Bera 等AAAI 2020 · 被引用 282 次
- DeeperForensics-1.0: A Large-Scale Dataset for Real-World Face Forgery DetectionLiming Jiang, Ren Li, Wayne Wu, Chen Qian 等CVPR 2020
- STEP: Spatial Temporal Graph Convolutional Networks for Emotion Perception from GaitsUttaran Bhattacharya, Trisha Mittal, Rohan Chandra, Tanmay Randhavane 等AAAI 2020
相关 Paper
- AVFF: Audio-Visual Feature Fusion for Video Deepfake DetectionTrevine Oorloff, Surya Koppisetti, Nicolò Bonettini, Divyaraj Solanki 等CVPR 2024 · 被引用 51 次
- Multi-modal Deepfake Detection via Multi-task Audio-Visual Prompt LearningHui Miao, Yuanfang Guo, Zeming Liu, Yunhong WangAAAI 2025 · 被引用 8 次
- Joint Audio-Visual Deepfake DetectionYipin Zhou, Ser-Nam LimICCV 2021 · 被引用 232 次
- SepVAMark: Deep Separable Visual-Audio Fusion Watermarking for Source Tracing and Deepfake DetectionChuan Zhang, Zihan Li, Zihao Xu, Xuhao Ren 等ACM MM 2025 · 被引用 2 次
- SafeEar: Content Privacy-Preserving Audio Deepfake DetectionXinfeng Li, Kai Li, Yifan Zheng, Chen Yan 等CCS 2024 · 被引用 26 次
