"Hello, It's Me": Deep Learning-based Speech Synthesis Attacks in the Real World
Emily Wenger, Max Bronckers, Christian Cianfarani, Jenna Cryan, Angela Sha, Haitao Zheng, Ben Y. Zhao
Abstract
Advances in deep learning have introduced a new wave of voice synthesis tools, capable of producing audio that sounds as if spoken by a target speaker. If successful, such tools in the wrong hands will enable a range of powerful attacks against both humans and software systems (aka machines). This paper documents efforts and findings from a comprehensive experimental study on the impact of deep-learning based speech synthesis attacks on both human listeners and machines such as speaker recognition and voice-signin systems. We find that both humans and machines can be reliably fooled by synthetic speech, and that existing defenses against synthesized speech fall short. These findings highlight the need to raise awareness and develop new protections against synthetic speech for both humans and machines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- AntiFake: Using Adversarial Audio to Prevent Unauthorized Speech SynthesisZhiyuan Yu, Shixuan Zhai, Ning ZhangCCS 2023 · 29 citations
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
- "Better Be Computer or I'm Dumb": A Large-Scale Evaluation of Humans as Audio Deepfake DetectorsKevin Warren, Tyler Tucker, Anna Crowder, Daniel Olszewski et al.CCS 2024 · 9 citations
- Perception-Aware Attack: Creating Adversarial Music via Reverse-Engineering Human PerceptionRui Duan, Zhe Qu, Shangqing Zhao, Leah Ding et al.CCS 2022 · 8 citations
- Can I Hear Your Face? Pervasive Attack on Voice Authentication Systems with a Single Face ImageNan Jiang, Bangjie Sun, Terence Sim, Jun HanUSENIX Security 2024 · 7 citations
Builds on7
- Who is Real Bob? Adversarial Attacks on Speaker Recognition SystemsGuangke Chen, Sen Chen, Lingling Fan, Xiaoning Du et al.S&P 2021 · 239 citations
- Hearing Your Voice is Not Enough: An Articulatory Gesture Based Liveness Detection for Voice AuthenticationLinghan Zhang, Sheng Tan, Jie YangCCS 2017 · 212 citations
- Unsupervised Speech Decomposition via Triple Information BottleneckKaizhi Qian, Yang Zhang, Shiyu Chang, Mark Hasegawa-Johnson et al.ICML 2020 · 210 citations
- DeepSonar: Towards Effective and Robust Detection of AI-Synthesized Fake VoicesRun Wang, Felix Juefei-Xu, Yihao Huang, Qing Guo et al.ACM MM 2020 · 124 citations
- The Catcher in the Field: A Fieldprint based Spoofing Detection for Text-Independent Speaker VerificationChen Yan, Yan Long, Xiaoyu Ji, Wenyuan XuCCS 2019 · 62 citations
Related papers
- PhoneyTalker: An Out-of-the-Box Toolkit for Adversarial Example Attack on Speaker RecognitionMeng Chen, Li Lu, Zhongjie Ba, Kui RenINFOCOM 2022 · 13 citations
- Who Are You (I Really Wanna Know)? Detecting Audio DeepFakes Through Vocal Tract ReconstructionLogan Blue, Kevin Warren, Hadi Abdullah, Cassidy Gibson et al.USENIX Security 2022
- Synthesising Audio Adversarial Examples for Automatic Speech RecognitionXinghua Qu, Pengfei Wei, Mingyong Gao, Zhu Sun et al.KDD 2022 · 7 citations
- VoiceRadar: Voice Deepfake Detection using Micro-Frequency and Compositional AnalysisKavita Kumari, Maryam Abbasihafshejani, Alessandro Pegoraro, Phillip Rieger et al.NDSS 2025
- SoK: The Faults in our ASRs: An Overview of Attacks against Automatic Speech Recognition and Speaker Identification SystemsHadi Abdullah, Kevin Warren, Vincent Bindschaedler, Nicolas Papernot et al.S&P 2021 · 145 citations
