Ok Google, What Am I Doing?: Acoustic Activity Recognition Bounded by Conversational Assistant Interactions
Rebecca Adaimi, Howard Yong, Edison Thomaz
Abstract
Conversational assistants in the form of stand-alone devices such as Amazon Echo and Google Home have become popular and embraced by millions of people. By serving as a natural interface to services ranging from home automation to media players, conversational assistants help people perform many tasks with ease, such as setting timers, playing music and managing to-do lists. While these systems offer useful capabilities, they are largely passive and unaware of the human behavioral context in which they are used. In this work, we explore how off-the-shelf conversational assistants can be enhanced with acoustic-based human activity recognition by leveraging the short interval after a voice command is given to the device. Since always-on audio recording can pose privacy concerns, our method is unique in that it does not require capturing and analyzing any audio other than the speech-based interactions between people and their conversational assistants. In particular, we leverage background environmental sounds present in these short duration voice-based interactions to recognize activities of daily living. We conducted a study with 14 participants in 3 different locations in their own homes. We showed that our method can recognize 19 different activities of daily living with average precision of 84.85% and average recall of 85.67% in a leave-one-participant-out performance evaluation with 30-second audio clips bound by the voice interactions.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1736c2ad-6d28-4989-b897-b7dff7374b8dCited by top-tier papers6
- Cosmo: contrastive fusion learning with small data for multimodal human activity recognitionXiaomin Ouyang, Xian Shuai, Jiayu Zhou, Ivy Wang Shi et al.MobiCom 2022 · 94 citations
- Understanding User Perceptions of Proactive Smart SpeakersJing Wei, Tilman Dingler, Vassilis KostakosUbiComp 2022 · 33 citations
- Automated Face-To-Face Conversation Detection on a Commodity Smartwatch with Acoustic SensingDawei Liang, Alice Zhang, Edison ThomazUbiComp 2023 · 11 citations
- ExpresSense: Exploring a Standalone Smartphone to Sense Engagement of Users from Facial Expressions Using Acoustic SensingPragma Kar, Shyamvanshikumar Singh, Avijit Mandal, Samiran Chattopadhyay et al.CHI 2023 · 7 citations
- Exploring Cultural and Intergenerational Dynamics in Voice Assistant Design for Chinese Older AdultsZhigu Qian, Jiaojiao Fu, Yangfan ZhouUbiComp 2025 · 6 citations
Related papers
- Hello There! Is Now a Good Time to Talk?: Opportune Moments for Proactive Interactions with Smart SpeakersNarae Cha, Auk Kim, Cheul Young Park, Soowon Kang et al.UbiComp 2020 · 74 citations
- EchoLIFE: Zero-Shot In-Home ADL Recognition with LLM-Guided Active Acoustic SensingYubin Lan, Qian Zhang, Shukai Ma, Changfei Dong et al.UbiComp 2026
- Investigating Users' Preferences and Expectations for Always-Listening Voice AssistantsMadiha Tabassum, Tomasz Kosinski, Alisa Frik, Nathan Malkin et al.UbiComp 2020 · 74 citations
- Leveraging Sound and Wrist Motion to Detect Activities of Daily Living with Commodity SmartwatchesSarnab Bhattacharya, Rebecca Adaimi, Edison ThomazUbiComp 2022 · 41 citations
- Spying through Your Voice Assistants: Realistic Voice Command FingerprintingDilawer Ahmed, Aafaq Sabir, Anupam DasUSENIX Security 2023
