Attacks as Defenses: Designing Robust Audio CAPTCHAs Using Attacks on Automatic Speech Recognition Systems
Hadi Abdullah, Aditya Karlekar, Saurabh Prasad, Muhammad Sajidur Rahman, Logan Blue, Luke A. Bauer, Vincent Bindschaedler, Patrick Traynor
Abstract
Audio CAPTCHAs are supposed to provide a strong defense for online resources; however, advances in speech-to-text mechanisms have rendered these defenses ineffective. Audio CAPTCHAs cannot simply be abandoned, as they are specifically named by the W3C as important enablers of accessibility. Accordingly, demonstrably more robust audio CAPTCHAs are important to the future of a secure and accessible Web. We look to recent literature on attacks on speech-to-text systems for inspiration for the construction of robust, principle-driven audio defenses. We begin by comparing 20 recent attack papers, classifying and measuring their suitability to serve as the basis of new"robust to transcription"but"easy for humans to understand"CAPTCHAs. After showing that none of these attacks alone are sufficient, we propose a new mechanism that is both comparatively intelligible (evaluated through a user study) and hard to automatically transcribe (i.e., ). Finally, we demonstrate that our audio samples have a high probability of being detected as CAPTCHAs when given to speech-to-text systems (). In so doing, we not only demonstrate a CAPTCHA that is approximately four orders of magnitude more difficult to crack, but that such systems can be designed based on the insights gained from attack papers using the differences between the ways that humans and computers process audio.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 640c9d2c-29d9-43f0-b0d9-ff43a0ea19a7Cited by top-tier papers2
- Compensating Removed Frequency Components: Thwarting Voice Spectrum Reduction AttacksShu Wang, Kun Sun, Qi LiNDSS 2024
- Are CAPTCHAs Still Bot-hard? Generalized Visual CAPTCHA Solving with Agentic Vision Language ModelXiwen Teoh, Yun Lin, Siqi Li, Ruofan Liu et al.USENIX Security 2025
Builds on12
- Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionMahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. ReiterCCS 2016 · 1,765 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
- DolphinAttack: Inaudible Voice CommandsGuoming Zhang, Chen Yan, Xiaoyu Ji, Tianchen Zhang et al.CCS 2017 · 753 citations
- Hidden Voice CommandsNicholas Carlini, Pratyush Mishra, Tavish Vaidya, Yuankai Zhang et al.USENIX Security 2016 · 672 citations
- CommanderSong: A Systematic Approach for Practical Adversarial Voice RecognitionXuejing Yuan, Yuxuan Chen, Yue Zhao, Yunhui Long et al.USENIX Security 2018 · 389 citations
Related papers
- A Simple Generic Attack on Text CaptchasHaichang Gao, Jeff Yan, Fang Cao, Zhengya Zhang et al.NDSS 2016 · 106 citations
- Text Captcha Is Dead? A Large Scale Deployment and Empirical StudyChenghui Shi, Shouling Ji, Qianjun Liu, Changchang Liu et al.CCS 2020 · 22 citations
- WaveGuard: Understanding and Mitigating Audio Adversarial ExamplesShehzeen Hussain, Paarth Neekhara, Shlomo Dubnov, Julian J. McAuley et al.USENIX Security 2021 · 89 citations
- IllusionCAPTCHA: A CAPTCHA based on Visual IllusionZiqi Ding, Gelei Deng, Yi Liu, Junchen Ding et al.WWW 2025 · 15 citations
- WavePurifier: Purifying Audio Adversarial Examples via Hierarchical Diffusion ModelsHanqing Guo, Guangjing Wang, Bocheng Chen, Yuanda Wang et al.MobiCom 2024 · 3 citations
