What You Read Isn't What You Hear: Linguistic Sensitivity in Deepfake Speech Detection
Binh Nguyen, Shuju Shi, Ryan Ofman, Thai Le
Abstract
Recent advances in text-to-speech technology have enabled highly realistic voice generation, fueling audio-based deepfake attacks such as fraud and impersonation. While audio antispoofing systems are critical for detecting such threats, prior research has predominantly focused on acoustic-level perturbations, leaving the impact of linguistic variation largely unexplored. In this paper, we investigate the linguistic sensitivity of both open-source and commercial anti-spoofing detectors by introducing TAPAS (Transcript-to-Audio Perturbation Anti-Spoofing), a novel framework for transcript-level adversarial attacks. Our extensive evaluation shows that even minor linguistic perturbations can significantly degrade detection accuracy: attack success rates exceed 60% on several open-source detector-voice pairs, and the accuracy of one commercial detector drops from 100% on synthetic audio to just 32%. Through a comprehensive feature attribution analysis, we find that linguistic complexity and model-level audio embedding similarity are key factors contributing to detector vulnerabilities. To illustrate the real-world risks, we replicate a recent Brad Pitt audio deepfake scam and demonstrate that TAPAS can bypass commercial detectors. These findings underscore the need to move beyond purely acoustic defenses and incorporate linguistic variation into the design of robust anti-spoofing systems. Our source code is available at https: //github.com/nqbinh17/audio_linguist ic_adversarial.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext acc6e7e4-206c-44da-9f64-9211392bf7efBuilds on5
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue et al.EMNLP 2020 · 529 citations
- Transferring Audio Deepfake Detection Capability across LanguagesZhongjie Ba, Qing Wen, Peng Cheng, Yuwei Wang et al.WWW 2023 · 34 citations
- VoiceWukong: Benchmarking Deepfake Voice DetectionZiwei Yan, Yanjie Zhao, Haoyu WangUSENIX Security 2025
- F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow MatchingYushen Chen, Zhikang Niu, Ziyang Ma, Keqi Deng et al.ACL 2025
Related papers
- "Better Be Computer or I'm Dumb": A Large-Scale Evaluation of Humans as Audio Deepfake DetectorsKevin Warren, Tyler Tucker, Anna Crowder, Daniel Olszewski et al.CCS 2024 · 9 citations
- AVA: Inconspicuous Attribute Variation-based Adversarial Attack bypassing DeepFake DetectionXiangtao Meng, Li Wang, Shanqing Guo, Lei Ju et al.S&P 2024 · 17 citations
- SiFDetectCracker: An Adversarial Attack Against Fake Voice Detection Based on Speaker-Irrelative FeaturesXuan Hai, Xin Liu, Yuan Tan, Qingguo ZhouACM MM 2023 · 6 citations
- Deepfake Text Detection: Limitations and OpportunitiesJiameng Pu, Zain Sarwar, Sifat Muhammad Abdullah, Abdullah Rehman et al.S&P 2023
- SafeEar: Content Privacy-Preserving Audio Deepfake DetectionXinfeng Li, Kai Li, Yifan Zheng, Chen Yan et al.CCS 2024 · 26 citations
