When Evil Calls: Targeted Adversarial Voice over IP Network
Han Liu, Zhiyuan Yu, Mingming Zha, XiaoFeng Wang, William Yeoh, Yevgeniy Vorobeychik, Ning Zhang
Abstract
As the COVID-19 pandemic fundamentally reshaped the remote life and working styles, Voice over IP (VoIP) telephony and video conferencing have become a primary method of connecting communities together. However, little has been done to understand the feasibility and limitations of delivering adversarial voice samples via such communication channels. In this paper, we propose TAINT -Targeted Adversarial Voice over IP Network, the first targeted, query-efficient, hard label blackbox, adversarial attack on commercial speech recognition platforms over VoIP. The unique channel characteristics of VoIP pose significant new challenges, such as signal degradation, random channel noise, frequency selectivity, etc. To address these challenges, we systematically analyze the structure and channel characteristics of VoIP through reverse engineering. A noise-resilient efficient gradient estimation method is then developed to ensure a steady and fast convergence of the adversarial sample generation process. We demonstrate our attack in both over-the-air and over-the-line settings on four commercial automatic speech recognition (ASR) systems over the five most popular VoIP Conferencing Software (VCS). We show that TAINT can achieve performance that is comparable to the existing methods even with the addition of VoIP channel. Even in the most challenging scenario where there is an active speaker in Zoom, TAINT can still succeed within 10 attempts while staying out of the speaker focus of the video conference 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ed34a24f-2ea3-48aa-b595-9b7702fa6567Cited by top-tier papers7
- ALIF: Low-Cost Adversarial Audio Attacks on Black-Box Speech Platforms using Linguistic FeaturesPeng Cheng, Yuwei Wang, Peng Huang, Zhongjie Ba et al.S&P 2024 · 16 citations
- AttackGNN: Red-Teaming GNNs in Hardware Security Using Reinforcement LearningVasudev Gohil, Satwik Patnaik, Dileep Kalathil, Jeyavijayan RajendranUSENIX Security 2024 · 9 citations
- IP Protection in TinyMLJinwen Wang, Yuhao Wu, Han Liu, Bo Yuan et al.DAC 2023 · 6 citations
- RIATIG: Reliable and Imperceptible Adversarial Text-to-Image Generation with Natural PromptsHan Liu, Yuhao Wu, Shixuan Zhai, Bo Yuan et al.CVPR 2023
- Whispering Under the Eaves: Protecting User Privacy Against Commercial and LLM-powered Automatic Speech Recognition SystemsWeifei Jin, Yuxin Cao, Junjie Su, Derui Wang et al.USENIX Security 2025
Builds on12
- DolphinAttack: Inaudible Voice CommandsGuoming Zhang, Chen Yan, Xiaoyu Ji, Tianchen Zhang et al.CCS 2017 · 753 citations
- Hidden Voice CommandsNicholas Carlini, Pratyush Mishra, Tavish Vaidya, Yuankai Zhang et al.USENIX Security 2016 · 672 citations
- DDSP: Differentiable Digital Signal ProcessingJesse H. Engel, Lamtharn Hantrakul, Chenjie Gu, Adam RobertsICLR 2020 · 467 citations
- CommanderSong: A Systematic Approach for Practical Adversarial Voice RecognitionXuejing Yuan, Yuxuan Chen, Yue Zhao, Yunhui Long et al.USENIX Security 2018 · 389 citations
- Adversarial Attacks Against Automatic Speech Recognition Systems via Psychoacoustic HidingLea Schönherr, Katharina Kohls, Steffen Zeiler, Thorsten Holz et al.NDSS 2019 · 315 citations
Related papers
- Devil's Whisper: A General Approach for Physical Adversarial Attacks against Commercial Black-box Speech Recognition DevicesYuxuan Chen, Xuejing Yuan, Jiangshan Zhang, Yue Zhao et al.USENIX Security 2020
- Zero-Query Adversarial Attack on Black-box Automatic Speech Recognition SystemsZheng Fang, Tao Wang, Lingchen Zhao, Shenyi Zhang et al.CCS 2024 · 11 citations
- Echo: Reverberation-based Fast Black-Box Adversarial Attacks on Intelligent Audio SystemsMeng Xue, Kuang Peng, Xueluan Gong, Qian Zhang et al.UbiComp 2023 · 2 citations
- KENKU: Towards Efficient and Stealthy Black-box Adversarial Attacks against ASR SystemsXinghui Wu, Shiqing Ma, Chao Shen, Chenhao Lin et al.USENIX Security 2023
- More Simplicity for Trainers, More Opportunity for Attackers: Black-Box Attacks on Speaker Recognition Systems by Inferring Feature ExtractorYunjie Ge, Pinji Chen, Qian Wang, Lingchen Zhao et al.USENIX Security 2024 · 4 citations
