Privacy against Real-Time Speech Emotion Detection via Acoustic Adversarial Evasion of Machine Learning
Brian Testa, Yi Xiao, Harshit Sharma, Avery Gump, Asif Salekin
Abstract
Smart speaker voice assistants (VAs) such as Amazon Echo and Google Home have been widely adopted due to their seamless integration with smart home devices and the Internet of Things (IoT) technologies. These VA services raise privacy concerns, especially due to their access to our speech. This work considers one such use case: the unaccountable and unauthorized surveillance of a user's emotion via speech emotion recognition (SER). This paper presents DARE-GP, a solution that creates additive noise to mask users' emotional information while preserving the transcription-relevant portions of their speech. DARE-GP does this by using a constrained genetic programming approach to learn the spectral frequency traits that depict target users' emotional content, and then generating a universal adversarial audio perturbation that provides this privacy protection. Unlike existing works, DARE-GP provides: a) real-time protection of previously unheard utterances, b) against previously unseen black-box SER classifiers, c) while protecting speech transcription, and d) does so in a realistic, acoustic environment. Further, this evasion is robust against defenses employed by a knowledgeable adversary. The evaluations in this work culminate with acoustic evaluations against two off-the-shelf commercial smart speakers using a small-form-factor (raspberry pi) integrated with a wake-word system to evaluate the efficacy of its real-world, real-time deployment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b49ebaec-c805-409f-a4ff-6af8464c34cbBuilds on10
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- DolphinAttack: Inaudible Voice CommandsGuoming Zhang, Chen Yan, Xiaoyu Ji, Tianchen Zhang et al.CCS 2017 · 753 citations
- Why Do Adversarial Attacks Transfer? Explaining Transferability of Evasion and Poisoning AttacksAmbra Demontis, Marco Melis, Maura Pintor, Matthew Jagielski et al.USENIX Security 2019 · 466 citations
- Sign-OPT: A Query-Efficient Hard-label Adversarial AttackMinhao Cheng, Simranjit Singh, Patrick H. Chen, Pin-Yu Chen et al.ICLR 2020 · 256 citations
- Automatically Evading Classifiers: A Case Study on PDF Malware ClassifiersWeilin Xu, Yanjun Qi, David EvansNDSS 2016 · 249 citations
Related papers
- Whispering Under the Eaves: Protecting User Privacy Against Commercial and LLM-powered Automatic Speech Recognition SystemsWeifei Jin, Yuxin Cao, Junjie Su, Derui Wang et al.USENIX Security 2025
- SafeSpeaker: Voice Obfuscation for Resource-Constrained IoT DevicesCameron Haire, Yasha Iravantchi, Kang G. Shin, Alanson P. SampleUbiComp 2026 · 1 citation
- Dangerous Skills: Understanding and Mitigating Security Risks of Voice-Controlled Third-Party Functions on Virtual Personal Assistant SystemsNan Zhang, Xianghang Mi, Xuan Feng, XiaoFeng Wang et al.S&P 2019 · 160 citations
- Devil's Whisper: A General Approach for Physical Adversarial Attacks against Commercial Black-box Speech Recognition DevicesYuxuan Chen, Xuejing Yuan, Jiangshan Zhang, Yue Zhao et al.USENIX Security 2020
- FakeWake: Understanding and Mitigating Fake Wake-up Words of Voice AssistantsYanjiao Chen, Yijie Bai, Richard Mitev, Kaibo Wang et al.CCS 2021 · 24 citations
