Time Shuffle: A Transferability-Booster for Multiple Audio Adversarial Tasks
Jiacheng Deng, Dengpan Ye, Yuhong Liu, Zhaolin Wei, Ziyi Liu, Haoran Duan
摘要
Existing audio adversarial attack methods suffer from poor transferability, primarily due to insufficient exploration of model decision mechanisms and overreliance on heuristic-driven algorithm design. This paper aims to alleviate this gap. Specifically, through observations across three mainstream audio tasks (Automatic Speech Recognition, Speaker Verification, and Keyword Spotting), we reveal that these models primarily rely on local temporal features—inputs with time shuffled retain 83.7% of original accuracy. The SHAP-based visualization further validated that time shuffle leads to a significant shift in the salient regions of the model, but the samples can still be correctly identified, indicating the presence of redundant features that can affect decision-making. Inspired by these findings, we propose Time-Shuffle (TS) adversarial attack (including segments-based TS and phoneme-level-based TS-p). This method divides audio or phonemes into segments, randomly shuffles them, and computes gradients on the shuffled structure. By forcing perturbations to exploit transferable local temporal features and reduce overfitting to source-specific patterns, TS/TS-p inherently enhances transferability. As a model-agnostic framework, TS/TS-p can seamlessly integrate with existing attack methods. Comprehensive experiments demonstrate that TS-p achieved SOTA and boosts transferability by about 23%/14.7%/6.3% on ASR/ASV/KWS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Adv-Attribute: Inconspicuous and Transferable Adversarial Attack on Face RecognitionShuai Jia, Bangjie Yin, Taiping Yao, Shouhong Ding 等NeurIPS 2022 · 被引用 84 次
- Diversified Adversarial Attacks based on Conjugate Gradient MethodKeiichiro Yamamura, Haruki Sato, Nariaki Tateiwa, Nozomi Hata 等ICML 2022 · 被引用 18 次
- Zero-Query Adversarial Attack on Black-box Automatic Speech Recognition SystemsZheng Fang, Tao Wang, Lingchen Zhao, Shenyi Zhang 等CCS 2024 · 被引用 11 次
- Enhancing Adversarial Transferability with Adversarial Weight TuningJiahao Chen, Zhou Feng, Rui Zeng, Yuwen Pu 等AAAI 2025 · 被引用 11 次
相关 Paper
- Demystifying Limited Adversarial Transferability in Automatic Speech Recognition SystemsHadi Abdullah, Aditya Karlekar, Vincent Bindschaedler, Patrick TraynorICLR 2022 · 被引用 12 次
- Boosting the Transferability of Video Adversarial Examples via Temporal TranslationZhipeng Wei, Jingjing Chen, Zuxuan Wu, Yu-Gang JiangAAAI 2022 · 被引用 48 次
- Rethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality PerspectivesZeliang Zhang, Susan Liang, Daiki Shimada, Chenliang XuICLR 2025
- Global-Local Characteristic Excited Cross-Modal Attacks from Images to VideosRuikui Wang, Yuanfang Guo, Yunhong WangAAAI 2023 · 被引用 15 次
- Cross-modal and Cross-medium Adversarial Attack for AudioLiguo Zhang, Zilin Tian, Yunfei Long, Sizhao Li 等ACM MM 2023 · 被引用 1 次
