The Crux of Voice (In)Security: A Brain Study of Speaker Legitimacy Detection
Ajaya Neupane, Nitesh Saxena, Leanne M. Hirshfield, Sarah E. Bratt
Abstract
A new generation of scams has emerged that uses voice impersonation to obtain sensitive information, eavesdrop over voice calls and extort money from unsuspecting human users. Research demonstrates that users are fallible to voice impersonation attacks that exploit the current advancement in speech synthesis. In this paper, we set out to elicit a deeper understanding of such human-centered “voice hacking” based on a neuro-scientific methodology (thereby corroborating and expanding the traditional behavioral-only approach in significant ways). Specifically, we investigate the neural underpinnings of voice security through functional near-infrared spectroscopy (fNIRS), a cutting-edge neuroimaging technique, that captures neural signals in both temporal and spatial domains. We design and conduct an fNIRS study to pursue a thorough investigation of users’ mental processing related to speaker legitimacy detection – whether a voice sample is rendered by a target speaker, a different other human speaker or a synthesizer mimicking the speaker. We analyze the neural activity associated within this task as well as the brain areas that may control such activity. Our key insight is that there may be no statistically significant differences in the way the human brain processes the legitimate speakers vs. synthesized speakers, whereas clear differences are visible when encountering legitimate vs. different other human speakers. This finding may help to explain users’ susceptibility to synthesized attacks, as seen from the behavioral self-reported analysis. That is, the impersonated synthesized voices may seem indistinguishable from the real voices in terms of both behavioral and neural perspectives. In sharp contrast, prior studies showed subconscious neural differences in other real vs. fake artifacts (e.g., paintings and websites), despite users failing to note these differences behaviorally. Overall, our work dissects the fundamental neural patterns underlying voice-based insecurity and reveals users’ susceptibility to voice synthesis attacks at a biological level. We believe that this could be a significant insight for the security community suggesting that the human detection of voice synthesis attacks may not improve over time, especially given that voice synthesis techniques will likely continue to improve, calling for the design of careful machine-assisted techniques to help humans counter these attacks. *Work done while being a student at UAB
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a01e966-9993-4d92-8b95-aaad7b0aa282Cited by top-tier papers5
- "Hello, It's Me": Deep Learning-based Speech Synthesis Attacks in the Real WorldEmily Wenger, Max Bronckers, Christian Cianfarani, Jenna Cryan et al.CCS 2021 · 36 citations
- KinVoices: Using Voices of Friends and Family in Voice InterfacesSamantha W. T. Chan, Tamil Selvan Gunasekaran, Yun Suen Pai, Haimo Zhang et al.CSCW 2021 · 32 citations
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
- "Better Be Computer or I'm Dumb": A Large-Scale Evaluation of Humans as Audio Deepfake DetectorsKevin Warren, Tyler Tucker, Anna Crowder, Daniel Olszewski et al.CCS 2024 · 9 citations
- DeepPhish: Understanding User Trust Towards Artificially Generated Profiles in Online Social NetworksJaron Mink, Licheng Luo, Natã M. Barbosa, Olivia Figueira et al.USENIX Security 2022
Builds on1
Related papers
- From One Stolen Utterance: Assessing the Risks of Voice Cloning in the AIGC EraKun Wang, Meng Chen, Li Lu, Jingwen Feng et al.S&P 2025
- Dangerous Skills: Understanding and Mitigating Security Risks of Voice-Controlled Third-Party Functions on Virtual Personal Assistant SystemsNan Zhang, Xianghang Mi, Xuan Feng, XiaoFeng Wang et al.S&P 2019 · 160 citations
- SiFDetectCracker: An Adversarial Attack Against Fake Voice Detection Based on Speaker-Irrelative FeaturesXuan Hai, Xin Liu, Yuan Tan, Qingguo ZhouACM MM 2023 · 6 citations
- VoiceRadar: Voice Deepfake Detection using Micro-Frequency and Compositional AnalysisKavita Kumari, Maryam Abbasihafshejani, Alessandro Pegoraro, Phillip Rieger et al.NDSS 2025
- DeepSonar: Towards Effective and Robust Detection of AI-Synthesized Fake VoicesRun Wang, Felix Juefei-Xu, Yihao Huang, Qing Guo et al.ACM MM 2020 · 124 citations
