VIPER Strike: Defeating Visual Reasoning CAPTCHAs via Structured Vision–Language Inference
Minfeng Qi, Dongyang He, Qin Wang, Lefeng Zhang
摘要
Visual Reasoning CAPTCHAs (VRCs) combine visual scenes with natural-language queries that demand compositional inference over objects, attributes, and spatial relations. They are increasingly deployed as a primary defense against automated bots. Existing solvers fall into two paradigms: vision-centric, which rely on template-specific detectors but fail on novel layouts, and reasoning-centric, which leverage LLMs but struggle with fine-grained visual perception. Both lack the generality needed to handle heterogeneous VRC deployments. We present VIPER, a unified attack framework that integrates structured multi-object visual perception with adaptive LLM-based reasoning. VIPER parses visual layouts, grounds attributes to question semantics, and infers target coordinates within a modular pipeline. Evaluated on six major VRC providers (VTT, Geetest, NetEase, Dingxiang, Shumei, Xiaodun), VIPER achieves up to 93.2% success, approaching human-level performance across multiple benchmarks. Compared to prior solvers, GraphNet (83.2%), Oedipus (65.8%), and the Holistic approach (89.5%), VIPER consistently outperforms all baselines. The framework further maintains robustness across alternative LLM backbones (GPT, Grok, DeepSeek, Kimi), sustaining accuracy above 90%. To anticipate defense, we further introduce Template-Space Randomization (TSR), a lightweight strategy that perturbs linguistic templates without altering task semantics. TSR measurably reduces solver (i.e., attacker) performance. Our proposed design suggests directions for human-solvable but machine-resistant CAPTCHAs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA SolversJunyu Wang, Changjia Zhu, Yuanbo Zhou, Lingyao Li 等USENIX Security 2026 · 被引用 4 次
- Next-Gen CAPTCHAs: Leveraging the Cognitive Gap for Scalable and Diverse GUI-Agent DefenseJiacheng Liu, Yaxin Luo, Jiacheng Cui, Xinyi Shang 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- Yet Another Text Captcha Solver: A Generative Adversarial Network Based ApproachGuixin Ye, Zhanyong Tang, Dingyi Fang, Zhanxing Zhu 等CCS 2018 · 被引用 138 次
相关 Paper
- Research on the Security of Visual Reasoning CAPTCHAYipeng Gao, Haichang Gao, Sainan Luo, Yang Zi 等USENIX Security 2021 · 被引用 19 次
- Are CAPTCHAs Still Bot-hard? Generalized Visual CAPTCHA Solving with Agentic Vision Language ModelXiwen Teoh, Yun Lin, Siqi Li, Ruofan Liu 等USENIX Security 2025
- Spatial CAPTCHA: Generatively Benchmarking Spatial Reasoning for Human-Machine DifferentiationArina Kharlamova, Bowei He, Chen Ma, Xue LiuICLR 2026 · 被引用 3 次
- Oedipus: LLM-enchanced Reasoning CAPTCHA SolverGelei Deng, Haoran Ou, Yi Liu, Jie Zhang 等CCS 2025
- ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language ModelsJuntian Zhang, Song Jin, Chuanqi Cheng, Yuhan Liu 等ICLR 2026 · 被引用 7 次
