VoiceBlock: Privacy through Real-Time Adversarial Attacks with Audio-to-Audio Models
Patrick O'Reilly, Andreas Bugler, Keshav Bhandari, Max Morrison, Bryan Pardo
Abstract
As governments and corporations adopt deep learning systems to collect and analyze user-generated audio data, concerns about security and privacy natu-rally emerge in areas such as automatic speaker recognition. While audio adversarial examples offer one route to mislead or evade these invasive systems, they are typically crafted through time-intensive offline optimization, limiting their usefulness in streaming contexts. Inspired by architectures for audio-to-audio tasks such as denoising and speech enhancement, we propose a neural network model capable of adversarially modifying a user’s audio stream in real-time. Our model learns to apply a time-varying finite impulse response (FIR) filter to outgoing audio, allowing for effective and inconspicuous perturbations on a small fixed delay suitable for streaming tasks. We demonstrate our model is highly effective at de-identifying user speech from speaker recognition and able to transfer to an unseen recognition system. We conduct a perceptual study and find that our method produces perturbations significantly less perceptible than baseline anonymization methods, when controlling for effectiveness. Finally, we provide an implementation of our model capable of running in real-time on a single CPU thread. Audio examples and code can be found at https://interactiveaudiolab.github.io/project/voiceblock.html .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca6a4d3c-391d-4ba8-90a5-2de12bb6c778Cited by top-tier papers3
- LaserAdv: Laser Adversarial Attacks on Speech Recognition SystemsGuoming Zhang, Xiaohui Ma, Huiting Zhang, Zhijie Xiang et al.USENIX Security 2024 · 6 citations
- Analyzing Inference Privacy Risks Through Gradients In Machine LearningZhuohang Li, Andrew Lowy, Jing Liu, Toshiaki Koike-Akino et al.CCS 2024 · 5 citations
- EveGuard: Defeating Vibration-based Side-Channel Eavesdropping with Audio Adversarial PerturbationsJung-Woo Chang, Ke Sun, David Xia, Xinyu Zhang et al.S&P 2025
Builds on7
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- DDSP: Differentiable Digital Signal ProcessingJesse H. Engel, Lamtharn Hantrakul, Chenjie Gu, Adam RobertsICLR 2020 · 467 citations
- AdvPulse: Universal, Synchronization-free, and Targeted Audio Adversarial Attacks via Subsecond PerturbationsZhuohang Li, Yi Wu, Jian Liu, Yingying Chen et al.CCS 2020 · 107 citations
- Enabling Fast and Universal Audio Adversarial Attack Using Generative ModelYi Xie, Zhuohang Li, Cong Shi, Jian Liu et al.AAAI 2021 · 77 citations
- Learning Transferable Adversarial PerturbationsKrishna Kanth Nakka, Mathieu SalzmannNeurIPS 2021 · 75 citations
Related papers
- Whispering Under the Eaves: Protecting User Privacy Against Commercial and LLM-powered Automatic Speech Recognition SystemsWeifei Jin, Yuxin Cao, Junjie Su, Derui Wang et al.USENIX Security 2025
- Real-Time Neural Voice CamouflageMia Chiquier, Chengzhi Mao, Carl VondrickICLR 2022 · 9 citations
- PhoneyTalker: An Out-of-the-Box Toolkit for Adversarial Example Attack on Speaker RecognitionMeng Chen, Li Lu, Zhongjie Ba, Kui RenINFOCOM 2022 · 13 citations
- VoiceCloak: Adversarial Example Enabled Voice De-Identification with Balanced Privacy and UtilityMeng Chen, Li Lu, Junhao Wang, Jiadi Yu et al.UbiComp 2023 · 21 citations
- V-Cloak: Intelligibility-, Naturalness- & Timbre-Preserving Real-Time Voice AnonymizationJiangyi Deng, Fei Teng, Yanjiao Chen, Xiaofu Chen et al.USENIX Security 2023
