Aware: Intuitive Device Activation Using Prosody for Natural Voice Interactions
Xinlei Zhang, Zixiong Su, Jun Rekimoto
Abstract
Voice interactive devices often use keyword spotting for device activation. However, this approach suffers from misrecognition of keywords and can respond to keywords not intended for calling the device (e.g., ”You can ask Alexa about it.”), causing accidental device activations. We propose a method that leverages prosodic features to differentiate calling/not-calling voices (F1 score: 0.869), allowing devices to respond only when called upon to avoid misactivation. As a proof of concept, we built a prototype smart speaker called Aware that allows users to control the device activation by speaking the keyword in specific prosody patterns. These patterns are chosen to represent people’s natural calling/not-calling voices, which are uncovered in a study to collect such voices and investigate their prosodic difference. A user study comparing Aware with Amazon Echo shows Aware can activate more correctly (F1 score 0.93 vs. 0.56) and is easy to learn and use.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 87fd4de6-286b-4cfd-b001-2c2c876dd3e5Related papers
- Towards More Robust Keyword Spotting for Voice AssistantsShimaa Ahmed, Ilia Shumailov, Nicolas Papernot, Kassem FawazUSENIX Security 2022
- Deaf and Hard of Hearing Access to Intelligent Personal Assistants: Comparison of Voice-Based Options with an LLM-Powered Touch InterfacePaige S. DeVries, Michaela Okosi, Ming Li, Nora Dunphy et al.CHI 2026 · 2 citations
- Ok Google, What Am I Doing?: Acoustic Activity Recognition Bounded by Conversational Assistant InteractionsRebecca Adaimi, Howard Yong, Edison ThomazUbiComp 2021 · 29 citations
- Understanding User Perceptions of Proactive Smart SpeakersJing Wei, Tilman Dingler, Vassilis KostakosUbiComp 2022 · 33 citations
- Direction-of-Voice (DoV) Estimation for Intuitive Speech Interaction with Smart Devices EcosystemsKaran Ahuja, Andy Kong, Mayank Goel, Chris HarrisonUIST 2020 · 29 citations
