Circumventing Shortcuts in Audio-visual Deepfake Detection Datasets with Unsupervised Learning
Stefan Smeu, Dragos-Alexandru Boldisor, Dan Oneata, Elisabeta Oneata
Abstract
Good datasets are essential for developing and benchmarking any machine learning system. Their importance is even more extreme for safety critical applications such as deepfake detection-the focus of this paper. Here we reveal that two of the most widely used audio-video deepfake datasets suffer from a previously unidentified spurious feature: the leading silence. Fake videos start with a very brief moment of silence and, on the basis of this feature alone, we can separate the real and fake samples almost perfectly. As such, previous audio-only and audio-video models exploit the presence of silence in the fake videos and consequently perform worse when the leading silence is removed. To circumvent latching on such an unwanted artifact and possibly other unrevealed ones, we propose a shift from supervised to unsupervised learning by training models exclusively on real data. We show that by aligning selfsupervised audio-video representations we remove the risk of relying on dataset-specific biases and improve robustness in deepfake detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d82faa5-849c-40a1-a9ce-08ec9ff0b84cCited by top-tier papers10
- TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake DetectionJian-Yu Jiang-Lin, Kang-Yang Huang, Ling Zou, Ling Lo et al.CVPR 2026 · 5 citations
- Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake DetectionTianxiao Li, Zhenglin Huang, Haiquan Wen, Yiwei He et al.CVPR 2026 · 5 citations
- X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake DetectionYoungseo Kim, Kwan Yun, Seokhyeon Hong, Sihun Cha et al.CVPR 2026 · 2 citations
- Investigating Self-Supervised Representations for Audio-Visual Deepfake DetectionDragos-Alexandru Boldisor, Stefan Smeu, Dan Oneata, Elisabeta OneataCVPR 2026 · 2 citations
- AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMsShuhan Xia, Peipei Li, Xuannan Liu, Dongsen Zhang et al.CVPR 2026 · 1 citation
Builds on22
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess et al.ICCV 2019 · 2,966 citations
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 1,267 citations
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- FSGAN: Subject Agnostic Face Swapping and ReenactmentYuval Nirkin, Yosi Keller, Tal HassnerICCV 2019 · 710 citations
- YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for EveryoneEdresson Casanova, Julian Weber, Christopher Dane Shulby, Arnaldo Cândido Júnior et al.ICML 2022 · 602 citations
Related papers
- Joint Audio-Visual Deepfake DetectionYipin Zhou, Ser-Nam LimICCV 2021 · 232 citations
- AVFF: Audio-Visual Feature Fusion for Video Deepfake DetectionTrevine Oorloff, Surya Koppisetti, Nicolò Bonettini, Divyaraj Solanki et al.CVPR 2024 · 51 citations
- SLIM: Style-Linguistics Mismatch Model for Generalized Audio Deepfake DetectionYi Zhu, Surya Koppisetti, Trang Tran, Gaurav BharajNeurIPS 2024 · 41 citations
- Multi-modal Deepfake Detection via Multi-task Audio-Visual Prompt LearningHui Miao, Yuanfang Guo, Zeming Liu, Yunhong WangAAAI 2025 · 8 citations
- AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake DatasetZhixi Cai, Shreya Ghosh, Aman Pankaj Adatia, Munawar Hayat et al.ACM MM 2024 · 51 citations
