Yet Another Text Captcha Solver: A Generative Adversarial Network Based Approach
Guixin Ye, Zhanyong Tang, Dingyi Fang, Zhanxing Zhu, Yansong Feng, Pengfei Xu, Xiaojiang Chen, Zheng Wang
Abstract
Despite several attacks have been proposed, text-based CAPTCHAs 1 are still being widely used as a security mechanism. One of the reasons for the pervasive use of text captchas is that many of the prior attacks are scheme-specific and require a labor-intensive and time-consuming process to construct. This means that a change in the captcha security features like a noisier background can simply invalid an earlier attack. This paper presents a generic, yet effective text captcha solver based on the generative adversarial network. Unlike prior machine-learning-based approaches that need a large volume of manually-labeled real captchas to learn an effective solver, our approach requires significantly fewer real captchas but yields much better performance. This is achieved by first learning a captcha synthesizer to automatically generate synthetic captchas to learn a base solver, and then fine-tuning the base solver on a small set of real captchas using transfer learning. We evaluate our approach by applying it to 33 captcha schemes, including 11 schemes that are currently being used by 32 of the top-50 popular websites including Microsoft, Wikipedia, eBay and Google. Our approach is the most capable attack on text captchas seen to date. It outperforms four state-of-the-art text-captcha solvers by not only delivering a significantly higher accuracy on all testing schemes, but also successfully attacking schemes where others have zero chance. We show that our approach is highly efficient as it can solve a captcha within 0.05 second using a desktop GPU. We demonstrate that our attack is generally applicable because it can bypass the advanced security features employed by most modern * Corresponding faculty authors: Zhanyong Tang and Zheng Wang. 1 To aid readability, we will use the acronym in lowercase thereafter.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d5b96469-cee7-47c5-bcbd-c9289bd6f2bfCited by top-tier papers18
- Throwing Darts in the Dark? Detecting Bots with Limited Data using Neural Data AugmentationSteve T. K. Jan, Qingying Hao, Tianrui Hu, Jiameng Pu et al.S&P 2020 · 88 citations
- Detecting Fake Accounts in Online Social Networks at the Time of RegistrationsDong Yuan, Yuanli Miao, Neil Zhenqiang Gong, Zheng Yang et al.CCS 2019 · 86 citations
- AdVersarial: Perceptual Ad Blocking meets Adversarial Machine LearningFlorian Tramèr, Pascal Dupré, Gili Rusak, Giancarlo Pellegrino et al.CCS 2019 · 65 citations
- Text Captcha Is Dead? A Large Scale Deployment and Empirical StudyChenghui Shi, Shouling Ji, Qianjun Liu, Changchang Liu et al.CCS 2020 · 22 citations
- Research on the Security of Visual Reasoning CAPTCHAYipeng Gao, Haichang Gao, Sainan Luo, Yang Zi et al.USENIX Security 2021 · 19 citations
Builds on2
Related papers
- A Generic Solver Combining Unsupervised Learning and Representation Learning for Breaking Text-Based CaptchasSheng Tian, Tao XiongWWW 2020 · 22 citations
- GeeSolver: A Generic, Efficient, and Effortless Solver with Self-Supervised Learning for Breaking Text CaptchasRuijie Zhao, Xianwen Deng, Yanhao Wang, Zhicong Yan et al.S&P 2023
- The Matter of Captchas: An Analysis of a Brittle Security Feature on the Modern WebBehzad Ousat, Esteban Schafir, Duc C. Hoang, Mohammad Ali Tofighi et al.WWW 2024 · 11 citations
- Are CAPTCHAs Still Bot-hard? Generalized Visual CAPTCHA Solving with Agentic Vision Language ModelXiwen Teoh, Yun Lin, Siqi Li, Ruofan Liu et al.USENIX Security 2025
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li et al.NDSS 2019 · 876 citations
