Alethia: a Foundational Encoder for Voice Deepfakes
Yi Zhu, Brahmi Dwivedi, Jayaram Raghuram, Surya Koppisetti
Abstract
Existing voice deepfake detection and localization models rely heavily on representations extracted from speech foundation models (SFMs). However, downstream finetuning has now reached a state of diminishing returns. In this paper, we shift the focus to pretraining and propose a novel recipe that combines bottleneck masked embedding prediction with flow-matching based spectrogram reconstruction . The outcome, Alethia , is the first foundational audio encoder for various voice deepfake detection and localization tasks. We evaluate on 5 different tasks with 56 benchmark datasets, and note Alethia significantly outperforms state-of-the-art SFMs with superior robustness to real-world perturbations and zero-shot generalization to unseen domains (e.g., singing deepfakes). We also demonstrate the limitation of discrete targets in masked token prediction, and show the importance of continuous embedding prediction and generative pretraining for capturing deepfake artifacts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3e3cba03-ab5e-4591-be30-e4028159a970Builds on8
- Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and LanguageAlexei Baevski, Arun Babu, Wei-Ning Hsu, Michael AuliICML 2023 · 137 citations
- Flow Matching for Generative ModelingYaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel et al.ICLR 2023 · 87 citations
- Generative Pre-training for Speech with Flow MatchingAlexander H. Liu, Matthew Le, Apoorv Vyas, Bowen Shi et al.ICLR 2024 · 66 citations
- Transferring Audio Deepfake Detection Capability across LanguagesZhongjie Ba, Qing Wen, Peng Cheng, Yuwei Wang et al.WWW 2023 · 34 citations
- Trident of Poseidon: A Generalized Approach for Detecting Deepfake VoicesThien-Phuc Doan, Hung Dinh-Xuan, Taewon Ryu, Inho Kim et al.CCS 2024 · 5 citations
Related papers
- Scaling Behavior in Model Fine-tuning for Audio DeepFake DetectionXiang Li, Pin-Yu Chen, Wenqi WeiICML 2026
- Investigating Self-Supervised Representations for Audio-Visual Deepfake DetectionDragos-Alexandru Boldisor, Stefan Smeu, Dan Oneata, Elisabeta OneataCVPR 2026 · 2 citations
- SLIM: Style-Linguistics Mismatch Model for Generalized Audio Deepfake DetectionYi Zhu, Surya Koppisetti, Trang Tran, Gaurav BharajNeurIPS 2024 · 41 citations
- Phoneme-Level Feature Discrepancies: A Key to Detecting Sophisticated Speech DeepfakesKuiyuan Zhang, Zhongyun Hua, Rushi Lan, Yushu Zhang et al.AAAI 2025 · 5 citations
- Circumventing Shortcuts in Audio-visual Deepfake Detection Datasets with Unsupervised LearningStefan Smeu, Dragos-Alexandru Boldisor, Dan Oneata, Elisabeta OneataCVPR 2025
