Kirigami: Lightweight Speech Filtering for Privacy-Preserving Activity Recognition using Audio
Sudershan Boovaraghavan, Haozhe Zhou, Mayank Goel, Yuvraj Agarwal
Abstract
Audio-based human activity recognition (HAR) is very popular because many human activities have unique sound signatures that can be detected using machine learning (ML) approaches. These audio-based ML HAR pipelines often use common featurization techniques, such as extracting various statistical and spectral features by converting time domain signals to the frequency domain (using an FFT) and using them to train ML models. Some of these approaches also claim privacy benefits by preventing the identification of human speech. However, recent deep learning-based automatic speech recognition (ASR) models pose new privacy challenges to these featurization techniques. In this paper, we systematically evaluate various featurization approaches for audio data, assessing their privacy risks through metrics like speech intelligibility (PER and WER) while considering the utility tradeoff in terms of ML-based activity recognition accuracy. Our findings reveal the susceptibility of these approaches to speech content recovery when exposed to recent ASR models, especially under re-tuning or retraining conditions. Notably, fine-tuned ASR models achieved an average Phoneme Error Rate (PER) of 39.99% and Word Error Rate (WER) of 44.43% in speech recognition for these approaches. To overcome these privacy concerns, we propose Kirigami, a lightweight machine learning-based audio speech filter that removes human speech segments reducing the efficacy of ASR models (70.48% PER and 101.40% WER) while also maintaining HAR accuracy (76.0% accuracy). We show that Kirigami can be implemented on common edge microcontrollers with limited computational capabilities and memory, providing a path to deployment on a variety of IoT devices. Finally, we conducted a real-world user study and showed the robustness of Kirigami on a laptop and an ARM Cortex-M4F microcontroller under three different background noises.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c53f5bff-1f1f-4e9a-9d62-ea484f6d0599Related papers
- PrivacyMic: Utilizing Inaudible Frequencies for Privacy Preserving Daily Activity RecognitionYasha Iravantchi, Karan Ahuja, Mayank Goel, Chris Harrison et al.CHI 2021 · 40 citations
- SAMoSA: Sensing Activities with Motion and Subsampled AudioVimal Mollyn, Karan Ahuja, Dhruv Verma, Chris Harrison et al.UbiComp 2022 · 54 citations
- SILENCE: Protecting privacy in offloaded speech understanding on resource-constrained devicesDongqi Cai, Shangguang Wang, Zeling Zhang, Felix Xiaozhu Lin et al.NeurIPS 2024 · 2 citations
- VAX: Using Existing Video and Audio-based Activity Recognition Models to Bootstrap Privacy-Sensitive SensorsPrasoon Patidar, Mayank Goel, Yuvraj AgarwalUbiComp 2023 · 12 citations
- SEGALL: A Unified Active Learning Framework for Wireless Sensing Data SegmentationNaiyu Zheng, Ruofeng Liu, Xiaoyi Fan, Cong Zhang et al.UbiComp 2025 · 3 citations
