SILENCE: Protecting privacy in offloaded speech understanding on resource-constrained devices
Dongqi Cai, Shangguang Wang, Zeling Zhang, Felix Xiaozhu Lin, Mengwei Xu
Abstract
Speech serves as a ubiquitous input interface for embedded mobile devices. Cloud-based solutions, while offering powerful speech understanding services, raise significant concerns regarding user privacy. To address this, disentanglement-based encoders have been proposed to remove sensitive information from speech signals without compromising the speech understanding functionality. However, these encoders demand high memory usage and computation complexity, making them impractical for resource-constrained wimpy devices. Our solution is based on a key observation that speech understanding hinges on long-term dependency knowledge of the entire utterance, in contrast to privacy-sensitive elements that are short-term dependent. Exploiting this observation, we propose SILENCE , a lightweight system that selectively obscuring short-term details, without damaging the long-term dependent speech understanding performance. The crucial part of SILENCE is a differential mask generator derived from interpretable learning to automatically configure the masking process. We have implemented SILENCE on the STM32H7 microcontroller and evaluate its efficacy under different attacking scenarios. Our results demonstrate that SILENCE offers speech understanding performance and privacy protection capacity comparable to existing encoders, while achieving up to 53.3 × speedup and 134.1 × reduction in memory footprint.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ce45f701-eb33-4c15-82e9-bd58abca3a17Builds on10
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- BatchCrypt: Efficient Homomorphic Encryption for Cross-Silo Federated LearningChengliang Zhang, Suyi Li, Junzhe Xia, Wei Wang et al.USENIX ATC 2020 · 967 citations
- Pengi: An Audio Language Model for Audio TasksSoham Deshmukh, Benjamin Elizalde, Rita Singh, Huaming WangNeurIPS 2023 · 352 citations
- SLURP: A Spoken Language Understanding Resource PackageEmanuele Bastianelli, Andrea Vanzo, Pawel Swietojanski, Verena RieserEMNLP 2020 · 129 citations
Related papers
- SafeSpeaker: Voice Obfuscation for Resource-Constrained IoT DevicesCameron Haire, Yasha Iravantchi, Kang G. Shin, Alanson P. SampleUbiComp 2026 · 1 citation
- Preech: A System for Privacy-Preserving Speech TranscriptionShimaa Ahmed, Amrita Roy Chowdhury, Kassem Fawaz, Parmesh RamanathanUSENIX Security 2020
- Knowing When to Quit: Probabilistic Early Exits for Speech Separation NetworksKenny Falkær Olsen, Mads Østergaard, Karl Ulbæk, Søren Føns Nielsen et al.ICLR 2026 · 1 citation
- MEPS: Privacy-preserving Edge-cloud Video Foundation Model Inference with Privacy ProtectabilitySiping Shi, Rui Lu, Dan Wang, Bihai ZhangUSENIX Security 2026
- Kirigami: Lightweight Speech Filtering for Privacy-Preserving Activity Recognition using AudioSudershan Boovaraghavan, Haozhe Zhou, Mayank Goel, Yuvraj AgarwalUbiComp 2024 · 7 citations
