Synthesising Audio Adversarial Examples for Automatic Speech Recognition
Xinghua Qu, Pengfei Wei, Mingyong Gao, Zhu Sun, Yew Soon Ong, Zejun Ma
摘要
Adversarial examples in automatic speech recognition (ASR) are naturally sounded by humans yet capable of fooling well trained ASR models to transcribe incorrectly. Existing audio adversarial examples are typically constructed by adding constrained perturbations on benign audio inputs. Such attacks are therefore generated with an audio dependent assumption. For the first time, we propose the Speech Synthesising based Attack (SSA), a novel threat model that constructs audio adversarial examples entirely from scratch, i.e., without depending on any existing audio to fool cutting-edge ASR models. To this end, we introduce a conditional variational auto-encoder (CVAE) as the speech synthesiser. Meanwhile, an adaptive sign gradient descent algorithm is proposed to solve the adversarial audio synthesis task. Experiments on three datasets (i.e., Audio Mnist, Common Voice, and Librispeech) show that our method could synthesise naturally sounded audio adversarial examples to mislead the start-of-the-art ASR models. Our web-page containing generated audio demos is at https://sites.google.com/view/ssa-asr/home.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- WaveGuard: Understanding and Mitigating Audio Adversarial ExamplesShehzeen Hussain, Paarth Neekhara, Shlomo Dubnov, Julian J. McAuley 等USENIX Security 2021 · 被引用 89 次
- A Unified Framework for Detecting Audio Adversarial ExamplesXia Du, Chi-Man Pun, Zheng ZhangACM MM 2020 · 被引用 18 次
- Weighted-Sampling Audio Adversarial Example AttackXiaolei Liu, Kun Wan, Yufei Ding, Xiaosong Zhang 等AAAI 2020 · 被引用 40 次
- Understanding and Benchmarking the Commonality of Adversarial ExamplesRuiwen He, Yushi Cheng, Junning Ze, Xiaoyu Ji 等S&P 2024 · 被引用 3 次
- Dompteur: Taming Audio Adversarial ExamplesThorsten Eisenhofer, Lea Schönherr, Joel Frank, Lars Speckemeier 等USENIX Security 2021 · 被引用 29 次
