DeepSonar: Towards Effective and Robust Detection of AI-Synthesized Fake Voices
Run Wang, Felix Juefei-Xu, Yihao Huang, Qing Guo, Xiaofei Xie, Lei Ma, Yang Liu
Abstract
With the recent advances in voice synthesis, AI-synthesized fake voices are indistinguishable to human ears and widely are applied to produce realistic and natural DeepFakes, exhibiting real threats to our society. However, effective and robust detectors for synthesized fake voices are still in their infancy and are not ready to fully tackle this emerging threat. In this paper, we devise a novel approach, named DeepSonar, based on monitoring neuron behaviors of speaker recognition (SR) system, i.e., a deep neural network (DNN), to discern AI-synthesized fake voices. Layer-wise neuron behaviors provide an important insight to meticulously catch the differences among inputs, which are widely employed for building safety, robust, and interpretable DNNs. In this work, we leverage the power of layer-wise neuron activation patterns with a conjecture that they can capture the subtle differences between real and AI-synthesized fake voices, in providing a cleaner signal to classifiers than raw inputs. Experiments are conducted on three datasets (including commercial products from Google, Baidu, etc.) containing both English and Chinese languages to corroborate the high detection rates (98.1% average accuracy) and low false alarm rates (about 2% error rate) of DeepSonar in discerning fake voices. Furthermore, extensive experimental results also demonstrate its robustness against manipulation attacks (e.g., voice conversion and additive real-world noises). Our work further poses a new insight into adopting neuron behaviors for effective and robust AI aided multimedia fakes forensics as an inside-out approach instead of being motivated and swayed by various artifacts introduced in synthesizing fakes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ebb06ddc-530a-4e93-bb10-00b1df2a6566Cited by top-tier papers15
- DeepRhythm: Exposing DeepFakes with Attentional Visual Heartbeat RhythmsHua Qi, Qing Guo, Felix Juefei-Xu, Xiaofei Xie et al.ACM MM 2020 · 224 citations
- Neural Analysis and Synthesis: Reconstructing Speech from Self-Supervised RepresentationsHyeong-Seok Choi, Juheon Lee, Wansoo Kim, Jie Lee et al.NeurIPS 2021 · 200 citations
- FakeTagger: Robust Safeguards against DeepFake Dissemination via Provenance TrackingRun Wang, Felix Juefei-Xu, Meng Luo, Yang Liu et al.ACM MM 2021 · 77 citations
- FakePolisher: Making DeepFakes More Detection-Evasive by Shallow ReconstructionYihao Huang, Felix Juefei-Xu, Run Wang, Qing Guo et al.ACM MM 2020 · 72 citations
- VoiceMixer: Adversarial Voice Style MixupSang-Hoon Lee, Ji-Hoon Kim, Hyunseung Chung, Seong-Whan LeeNeurIPS 2021 · 46 citations
Builds on6
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 840 citations
- NIC: Detecting Adversarial Samples with Neural Network Invariant CheckingShiqing Ma, Yingqi Liu, Guanhong Tao, Wen-Chuan Lee et al.NDSS 2019 · 283 citations
- DeepRhythm: Exposing DeepFakes with Attentional Visual Heartbeat RhythmsHua Qi, Qing Guo, Felix Juefei-Xu, Xiaofei Xie et al.ACM MM 2020 · 224 citations
- FakePolisher: Making DeepFakes More Detection-Evasive by Shallow ReconstructionYihao Huang, Felix Juefei-Xu, Run Wang, Qing Guo et al.ACM MM 2020 · 72 citations
- DeeperForensics-1.0: A Large-Scale Dataset for Real-World Face Forgery DetectionLiming Jiang, Ren Li, Wayne Wu, Chen Qian et al.CVPR 2020
Related papers
- SiFDetectCracker: An Adversarial Attack Against Fake Voice Detection Based on Speaker-Irrelative FeaturesXuan Hai, Xin Liu, Yuan Tan, Qingguo ZhouACM MM 2023 · 6 citations
- A Unified Framework for Detecting Audio Adversarial ExamplesXia Du, Chi-Man Pun, Zheng ZhangACM MM 2020 · 18 citations
- AntiFake: Using Adversarial Audio to Prevent Unauthorized Speech SynthesisZhiyuan Yu, Shixuan Zhai, Ning ZhangCCS 2023 · 29 citations
- Joint Audio-Visual Deepfake DetectionYipin Zhou, Ser-Nam LimICCV 2021 · 232 citations
- VoiceRadar: Voice Deepfake Detection using Micro-Frequency and Compositional AnalysisKavita Kumari, Maryam Abbasihafshejani, Alessandro Pegoraro, Phillip Rieger et al.NDSS 2025
