Alethia: a Foundational Encoder for Voice Deepfakes
Yi Zhu, Brahmi Dwivedi, Jayaram Raghuram, Surya Koppisetti
摘要
Existing voice deepfake detection and localization models rely heavily on representations extracted from speech foundation models (SFMs). However, downstream finetuning has now reached a state of diminishing returns. In this paper, we shift the focus to pretraining and propose a novel recipe that combines bottleneck masked embedding prediction with flow-matching based spectrogram reconstruction . The outcome, Alethia , is the first foundational audio encoder for various voice deepfake detection and localization tasks. We evaluate on 5 different tasks with 56 benchmark datasets, and note Alethia significantly outperforms state-of-the-art SFMs with superior robustness to real-world perturbations and zero-shot generalization to unseen domains (e.g., singing deepfakes). We also demonstrate the limitation of discrete targets in masked token prediction, and show the importance of continuous embedding prediction and generative pretraining for capturing deepfake artifacts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and LanguageAlexei Baevski, Arun Babu, Wei-Ning Hsu, Michael AuliICML 2023 · 被引用 137 次
- Flow Matching for Generative ModelingYaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel 等ICLR 2023 · 被引用 87 次
- Generative Pre-training for Speech with Flow MatchingAlexander H. Liu, Matthew Le, Apoorv Vyas, Bowen Shi 等ICLR 2024 · 被引用 66 次
- Transferring Audio Deepfake Detection Capability across LanguagesZhongjie Ba, Qing Wen, Peng Cheng, Yuwei Wang 等WWW 2023 · 被引用 34 次
- Trident of Poseidon: A Generalized Approach for Detecting Deepfake VoicesThien-Phuc Doan, Hung Dinh-Xuan, Taewon Ryu, Inho Kim 等CCS 2024 · 被引用 5 次
相关 Paper
- Scaling Behavior in Model Fine-tuning for Audio DeepFake DetectionXiang Li, Pin-Yu Chen, Wenqi WeiICML 2026
- Investigating Self-Supervised Representations for Audio-Visual Deepfake DetectionDragos-Alexandru Boldisor, Stefan Smeu, Dan Oneata, Elisabeta OneataCVPR 2026 · 被引用 2 次
- SLIM: Style-Linguistics Mismatch Model for Generalized Audio Deepfake DetectionYi Zhu, Surya Koppisetti, Trang Tran, Gaurav BharajNeurIPS 2024 · 被引用 41 次
- Phoneme-Level Feature Discrepancies: A Key to Detecting Sophisticated Speech DeepfakesKuiyuan Zhang, Zhongyun Hua, Rushi Lan, Yushu Zhang 等AAAI 2025 · 被引用 5 次
- Circumventing Shortcuts in Audio-visual Deepfake Detection Datasets with Unsupervised LearningStefan Smeu, Dragos-Alexandru Boldisor, Dan Oneata, Elisabeta OneataCVPR 2025
