Time Shuffle: A Transferability-Booster for Multiple Audio Adversarial Tasks
Jiacheng Deng, Dengpan Ye, Yuhong Liu, Zhaolin Wei, Ziyi Liu, Haoran Duan
Abstract
Existing audio adversarial attack methods suffer from poor transferability, primarily due to insufficient exploration of model decision mechanisms and overreliance on heuristic-driven algorithm design. This paper aims to alleviate this gap. Specifically, through observations across three mainstream audio tasks (Automatic Speech Recognition, Speaker Verification, and Keyword Spotting), we reveal that these models primarily rely on local temporal features—inputs with time shuffled retain 83.7% of original accuracy. The SHAP-based visualization further validated that time shuffle leads to a significant shift in the salient regions of the model, but the samples can still be correctly identified, indicating the presence of redundant features that can affect decision-making. Inspired by these findings, we propose Time-Shuffle (TS) adversarial attack (including segments-based TS and phoneme-level-based TS-p). This method divides audio or phonemes into segments, randomly shuffles them, and computes gradients on the shuffled structure. By forcing perturbations to exploit transferable local temporal features and reduce overfitting to source-specific patterns, TS/TS-p inherently enhances transferability. As a model-agnostic framework, TS/TS-p can seamlessly integrate with existing attack methods. Comprehensive experiments demonstrate that TS-p achieved SOTA and boosts transferability by about 23%/14.7%/6.3% on ASR/ASV/KWS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c2256adb-a021-4de2-83a8-00df468f32a5Builds on4
- Adv-Attribute: Inconspicuous and Transferable Adversarial Attack on Face RecognitionShuai Jia, Bangjie Yin, Taiping Yao, Shouhong Ding et al.NeurIPS 2022 · 84 citations
- Diversified Adversarial Attacks based on Conjugate Gradient MethodKeiichiro Yamamura, Haruki Sato, Nariaki Tateiwa, Nozomi Hata et al.ICML 2022 · 18 citations
- Zero-Query Adversarial Attack on Black-box Automatic Speech Recognition SystemsZheng Fang, Tao Wang, Lingchen Zhao, Shenyi Zhang et al.CCS 2024 · 11 citations
- Enhancing Adversarial Transferability with Adversarial Weight TuningJiahao Chen, Zhou Feng, Rui Zeng, Yuwen Pu et al.AAAI 2025 · 11 citations
Related papers
- Demystifying Limited Adversarial Transferability in Automatic Speech Recognition SystemsHadi Abdullah, Aditya Karlekar, Vincent Bindschaedler, Patrick TraynorICLR 2022 · 12 citations
- Boosting the Transferability of Video Adversarial Examples via Temporal TranslationZhipeng Wei, Jingjing Chen, Zuxuan Wu, Yu-Gang JiangAAAI 2022 · 48 citations
- Rethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality PerspectivesZeliang Zhang, Susan Liang, Daiki Shimada, Chenliang XuICLR 2025
- Global-Local Characteristic Excited Cross-Modal Attacks from Images to VideosRuikui Wang, Yuanfang Guo, Yunhong WangAAAI 2023 · 15 citations
- Cross-modal and Cross-medium Adversarial Attack for AudioLiguo Zhang, Zilin Tian, Yunfei Long, Sizhao Li et al.ACM MM 2023 · 1 citation
