USENIX Security2022Top-tier venue
Towards More Robust Keyword Spotting for Voice Assistants
Shimaa Ahmed, Ilia Shumailov, Nicolas Papernot, Kassem Fawaz
Abstract
Voice assistants rely on keyword spotting (KWS) to process vocal commands issued by humans: commands are prepended with a keyword, such as "Alexa" or "Ok Google," which must be spotted to activate the voice assistant. Typically, keyword spotting is two-fold: an on-device model first identifies the keyword, then the resulting voice sample triggers a second on-cloud model which verifies and processes the activation. In this work, we explore the significant privacy and security concerns that this raises under two threat models. First, our experiments demonstrate that accidental activations result in up to a minute of speech recording being uploaded to the cloud. Second, we verify that adversaries can systematically trigger misactivations through adversarial examples, which exposes the integrity and availability of services connected to the voice assistant. We propose EKOS (Ensemble for KeywOrd Spotting) which leverages the semantics of the KWS task to defend against both accidental and adversarial activations. EKOS incorporates spatial redundancy from the acoustic environment at training and inference time to minimize distribution drifts responsible for accidental activations. It also exploits a physical property of speech-its redundancy at different harmonics-to deploy an ensemble of models trained on different harmonics and provably force the adversary to modify more of the frequency spectrum to obtain adversarial examples. Our evaluation shows that EKOS increases the cost of adversarial activations, while preserving the natural accuracy. We validate the performance of EKOS with over-the-air experiments on commodity devices and commercial voice assistants; we find that EKOS improves the precision of the KWS task in non-adversarial settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e042cbe-ff91-4725-9916-38a04fc45e4cCited by top-tier papers8
- On the Limitations of Stochastic Pre-processing DefensesYue Gao, Ilia Shumailov, Kassem Fawaz, Nicolas PapernotNeurIPS 2022 · 35 citations
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
- Machine Learning needs Better Randomness Standards: Randomised Smoothing and PRNG-based attacksPranav Dahiya, Ilia Shumailov, Ross AndersonUSENIX Security 2024 · 11 citations
- SkillScanner: Detecting Policy-Violating Voice Applications Through Static Analysis at the Development PhaseSong Liao, Long Cheng, Haipeng Cai, Linke Guo et al.CCS 2023 · 7 citations
- From Virtual Touch to Tesla Command: Unlocking Unauthenticated Control Chains From Smart Glasses for Vehicle TakeoverXingli Zhang, Yazhou Tu, Yan Long, Liqun Shan et al.S&P 2024 · 6 citations
Builds on10
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos et al.USENIX Security 2019 · 1,386 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu et al.S&P 2019 · 1,022 citations
- DolphinAttack: Inaudible Voice CommandsGuoming Zhang, Chen Yan, Xiaoyu Ji, Tianchen Zhang et al.CCS 2017 · 753 citations
- Distance-Bounding Protocols: Verification without Time and LocationSjouke Mauw, Zach Smith, Jorge Toro-Pozo, Rolando Trujillo-RasuaS&P 2018 · 58 citations
Related papers
- Aware: Intuitive Device Activation Using Prosody for Natural Voice InteractionsXinlei Zhang, Zixiong Su, Jun RekimotoCHI 2022 · 7 citations
- Spying through Your Voice Assistants: Realistic Voice Command FingerprintingDilawer Ahmed, Aafaq Sabir, Anupam DasUSENIX Security 2023
- Dangerous Skills: Understanding and Mitigating Security Risks of Voice-Controlled Third-Party Functions on Virtual Personal Assistant SystemsNan Zhang, Xianghang Mi, Xuan Feng, XiaoFeng Wang et al.S&P 2019 · 160 citations
- FakeWake: Understanding and Mitigating Fake Wake-up Words of Voice AssistantsYanjiao Chen, Yijie Bai, Richard Mitev, Kaibo Wang et al.CCS 2021 · 24 citations
- VoiceCloak: Adversarial Example Enabled Voice De-Identification with Balanced Privacy and UtilityMeng Chen, Li Lu, Junhao Wang, Jiadi Yu et al.UbiComp 2023 · 21 citations
