SAMoSA: Sensing Activities with Motion and Subsampled Audio
Vimal Mollyn, Karan Ahuja, Dhruv Verma, Chris Harrison, Mayank Goel
Abstract
Despite advances in audio- and motion-based human activity recognition (HAR) systems, a practical, power-efficient, and privacy-sensitive activity recognition system has remained elusive. State-of-the-art activity recognition systems often require power-hungry and privacy-invasive audio data. This is especially challenging for resource-constrained wearables, such as smartwatches. To counter the need for an always-on audio-based activity classification system, we first make use of power and compute-optimized IMUs sampled at 50 Hz to act as a trigger for detecting activity events. Once detected, we use a multimodal deep learning model that augments the motion data with audio data captured on a smartwatch. We subsample this audio to rates ≤ 1 kHz, rendering spoken content unintelligible, while also reducing power consumption on mobile devices. Our multimodal deep learning model achieves a recognition accuracy of 92.2% across 26 daily activities in four indoor environments. Our findings show that subsampling audio from 16 kHz down to 1 kHz, in concert with motion data, does not result in a significant drop in inference accuracy. We also analyze the speech content intelligibility and power requirements of audio sampled at less than 1 kHz and demonstrate that our proposed approach can improve the practicality of human activity recognition systems.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get fb6999ce-2b69-440b-a816-ff05a124dd61Cited by top-tier papers15
- EchoWrist: Continuous Hand Pose Tracking and Hand-Object Interaction Recognition Using Low-Power Active Acoustic Sensing On a WristbandChi-Jung Lee, Ruidong Zhang, Devansh Agarwal, Tianhong Catherine Yu et al.CHI 2024 · 48 citations
- Semantic Hearing: Programming Acoustic Scenes with Binaural HearablesBandhav Veluri, Malek Itani, Justin Chan, Takuya Yoshioka et al.UIST 2023 · 29 citations
- ActSonic: Recognizing Everyday Activities from Inaudible Acoustic Wave Around the BodySaif Mahmud, Vineet Parikh, Qikang Liang, Ke Li et al.UbiComp 2025 · 24 citations
- PrISM-Tracker: A Framework for Multimodal Procedure Tracking Using Wearable Sensors and State Transition Information with User-Driven Handling of Errors and UncertaintyRiku Arakawa, Hiromu Yakura, Vimal Mollyn, Suzanne Nie et al.UbiComp 2023 · 20 citations
- PrISM-Q&A: Step-Aware Voice Assistant on a Smartwatch Enabled by Multimodal Procedure Tracking and Large Language ModelsRiku Arakawa, Jill Fain Lehman, Mayank GoelUbiComp 2025 · 20 citations
Related papers
- Leveraging Sound and Wrist Motion to Detect Activities of Daily Living with Commodity SmartwatchesSarnab Bhattacharya, Rebecca Adaimi, Edison ThomazUbiComp 2022 · 41 citations
- IMU2Doppler: Cross-Modal Domain Adaptation for Doppler-based Activity Recognition Using IMU DataSejal Bhalla, Mayank Goel, Rushil KhuranaUbiComp 2022 · 42 citations
- DeepSenseMoE: Harnessing Power of Time Series Foundation Models for Few-Shot Human Activity RecognitionZenan Fu, Dongzhou Cheng, Lei Zhang, Wenbo Huang et al.AAAI 2026
- Unsupervised Human Activity Representation Learning with Multi-task Deep ClusteringHaojie Ma, Zhijie Zhang, Wenzhong Li, Sanglu LuUbiComp 2021 · 46 citations
- Kirigami: Lightweight Speech Filtering for Privacy-Preserving Activity Recognition using AudioSudershan Boovaraghavan, Haozhe Zhou, Mayank Goel, Yuvraj AgarwalUbiComp 2024 · 7 citations
