COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers
Junyu Wang, Changjia Zhu, Yuanbo Zhou, Lingyao Li, Xu He, Mingkui Wei, Junjie Xiong
摘要
This paper studies how multimodal large language models (MLLMs) undermine the security guarantees of visual CAPTCHA. We identify the attack surface where an adversary can cheaply automate CAPTCHA solving using off-theshelf models. We evaluate 7 representative MLLMs on 18 real-world CAPTCHA task types, measuring single-shot accuracy, success under limited retries, end-to-end latency, and per-solve cost. We further validate our findings through a supplemental external dataset and an adaptive-attacker setting with session memory, while also analyzing the impact of taskspecific prompt engineering and few-shot demonstrations on solver effectiveness. We reveal that MLLMs can reliably solve recognition-oriented and low-interaction CAPTCHA tasks at human-like cost and latency, whereas tasks requiring finegrained localization, multi-step spatial reasoning, or crossframe consistency remain significantly harder for current models. By examining the reasoning traces of such MLLMs, we investigate the underlying mechanisms of why models succeed/fail on specific CAPTCHA puzzles and use these insights to derive defense-oriented guidelines for selecting and strengthening CAPTCHA tasks. To validate these principles, we present a proof-of-concept by hardening a vulnerable CAPTCHA type using our guidelines. We demonstrate that incorporating fine-grained localization and implicit counting reduces the success rate of state-of-the-art MLLMs from over 95% to 0%, confirming that structural changes can effectively mitigate the threat. We conclude by emphasizing the urgent need for CAPTCHA redesign as MLLM capabilities increasingly threaten existing defenses. Code Availability 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper12
- Yet Another Text Captcha Solver: A Generative Adversarial Network Based ApproachGuixin Ye, Zhanyong Tang, Dingyi Fang, Zhanxing Zhu 等CCS 2018 · 被引用 138 次
- A Generic Solver Combining Unsupervised Learning and Representation Learning for Breaking Text-Based CaptchasSheng Tian, Tao XiongWWW 2020 · 被引用 22 次
- Text Captcha Is Dead? A Large Scale Deployment and Empirical StudyChenghui Shi, Shouling Ji, Qianjun Liu, Changchang Liu 等CCS 2020 · 被引用 22 次
- Research on the Security of Visual Reasoning CAPTCHAYipeng Gao, Haichang Gao, Sainan Luo, Yang Zi 等USENIX Security 2021 · 被引用 19 次
- IllusionCAPTCHA: A CAPTCHA based on Visual IllusionZiqi Ding, Gelei Deng, Yi Liu, Junchen Ding 等WWW 2025 · 被引用 15 次
相关 Paper
- Oedipus: LLM-enchanced Reasoning CAPTCHA SolverGelei Deng, Haoran Ou, Yi Liu, Jie Zhang 等CCS 2025
- Spatial CAPTCHA: Generatively Benchmarking Spatial Reasoning for Human-Machine DifferentiationArina Kharlamova, Bowei He, Chen Ma, Xue LiuICLR 2026 · 被引用 3 次
- DualMirage: Hunting Stealthy Multimodal LLM Agents via CAPTCHAs with Contour and Adversarial IllusionsBei Chen, Gaolei Li, Jun Wu, Jianhua LiCVPR 2026
- MLLM-Protector: Ensuring MLLM's Safety without Hurting PerformanceRenjie Pi, Tianyang Han, Jianshu Zhang, Yueqi Xie 等EMNLP 2024 · 被引用 21 次
- Are CAPTCHAs Still Bot-hard? Generalized Visual CAPTCHA Solving with Agentic Vision Language ModelXiwen Teoh, Yun Lin, Siqi Li, Ruofan Liu 等USENIX Security 2025
