USENIX Security2026Top-tier venue
COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers
Junyu Wang, Changjia Zhu, Yuanbo Zhou, Lingyao Li, Xu He, Mingkui Wei, Junjie Xiong
Abstract
This paper studies how multimodal large language models (MLLMs) undermine the security guarantees of visual CAPTCHA. We identify the attack surface where an adversary can cheaply automate CAPTCHA solving using off-theshelf models. We evaluate 7 representative MLLMs on 18 real-world CAPTCHA task types, measuring single-shot accuracy, success under limited retries, end-to-end latency, and per-solve cost. We further validate our findings through a supplemental external dataset and an adaptive-attacker setting with session memory, while also analyzing the impact of taskspecific prompt engineering and few-shot demonstrations on solver effectiveness. We reveal that MLLMs can reliably solve recognition-oriented and low-interaction CAPTCHA tasks at human-like cost and latency, whereas tasks requiring finegrained localization, multi-step spatial reasoning, or crossframe consistency remain significantly harder for current models. By examining the reasoning traces of such MLLMs, we investigate the underlying mechanisms of why models succeed/fail on specific CAPTCHA puzzles and use these insights to derive defense-oriented guidelines for selecting and strengthening CAPTCHA tasks. To validate these principles, we present a proof-of-concept by hardening a vulnerable CAPTCHA type using our guidelines. We demonstrate that incorporating fine-grained localization and implicit counting reduces the success rate of state-of-the-art MLLMs from over 95% to 0%, confirming that structural changes can effectively mitigate the threat. We conclude by emphasizing the urgent need for CAPTCHA redesign as MLLM capabilities increasingly threaten existing defenses. Code Availability 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f311510-d0aa-4577-8dbd-e4f9aa69126aCited by top-tier papers1
Ask how each one uses itBuilds on12
- Yet Another Text Captcha Solver: A Generative Adversarial Network Based ApproachGuixin Ye, Zhanyong Tang, Dingyi Fang, Zhanxing Zhu et al.CCS 2018 · 138 citations
- A Generic Solver Combining Unsupervised Learning and Representation Learning for Breaking Text-Based CaptchasSheng Tian, Tao XiongWWW 2020 · 22 citations
- Text Captcha Is Dead? A Large Scale Deployment and Empirical StudyChenghui Shi, Shouling Ji, Qianjun Liu, Changchang Liu et al.CCS 2020 · 22 citations
- Research on the Security of Visual Reasoning CAPTCHAYipeng Gao, Haichang Gao, Sainan Luo, Yang Zi et al.USENIX Security 2021 · 19 citations
- IllusionCAPTCHA: A CAPTCHA based on Visual IllusionZiqi Ding, Gelei Deng, Yi Liu, Junchen Ding et al.WWW 2025 · 15 citations
Related papers
- Oedipus: LLM-enchanced Reasoning CAPTCHA SolverGelei Deng, Haoran Ou, Yi Liu, Jie Zhang et al.CCS 2025
- Spatial CAPTCHA: Generatively Benchmarking Spatial Reasoning for Human-Machine DifferentiationArina Kharlamova, Bowei He, Chen Ma, Xue LiuICLR 2026 · 3 citations
- DualMirage: Hunting Stealthy Multimodal LLM Agents via CAPTCHAs with Contour and Adversarial IllusionsBei Chen, Gaolei Li, Jun Wu, Jianhua LiCVPR 2026
- MLLM-Protector: Ensuring MLLM's Safety without Hurting PerformanceRenjie Pi, Tianyang Han, Jianshu Zhang, Yueqi Xie et al.EMNLP 2024 · 21 citations
- Are CAPTCHAs Still Bot-hard? Generalized Visual CAPTCHA Solving with Agentic Vision Language ModelXiwen Teoh, Yun Lin, Siqi Li, Ruofan Liu et al.USENIX Security 2025
