MirrorCAPTCHA: Wild CAPTCHA, Wild Distribution, Wild Web-based Platform Meet Multimodal LLM Agents
Xiangyu Wu, Yuwei Hu, Tianyu Cui, Yueying Tian, Qingguo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, Yang Yang, Jianfeng Lu
Abstract
The path to fully autonomous web agents is currently hindered by a critical bottleneck: their limited ability to handle CAPTCHA. Existing agent benchmarks largely ignore this practical challenge, failing to evaluate an agent’s real-world capacity to solve CAPTCHA. To bridge this gap, we conduct a comprehensive analysis of real-world CAPTCHA distributions and introduce MirrorCAPTCHA , a benchmark annotated with Weighted Pass Rate and a newly proposed metric Completion Degree . Mirror-CAPTCHA is designed to serve as a “mirror” that faithfully reflects the automation capabilities of agents in real scenarios. We filter 2 , 095 websites from Common Crawl, identify the CAPTCHA deployed on these sites, and cluster them into 18 distinct categories us-ing K-means algorithm. To ensure practicality, we extract a web subgraph from Common Crawl covering these websites and use random walks to simulate real-world CAPTCHA encounter frequencies, yielding a realistic measure of agents’ ability. Additionally, we develop a lightweight synthetic data pipeline to train Ovis2-Agent-CAPTCHA-8B , which significantly outperforms current state-of-the-art closed-source models on MirrorCAPTCHA, achieving a 9 . 4% higher average Weighted Pass Rate and a 2 . 13% higher average Completion Degree than the runner-up, Gemini-2.5-Pro .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- A Real-World WebAgent with Planning, Long Context Understanding, and Program SynthesisIzzeddin Gur, Hiroki Furuta, Austin V. Huang, Mustafa Safdari et al.ICLR 2024 · 359 citations
- ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web AgentsIdo Levy, Ben wiesel, Sami Marreed, Alon Oved et al.ICLR 2026 · 78 citations
- Division-of-Thoughts: Harnessing Hybrid Language Model Synergy for Efficient On-Device AgentsChenyang Shao, Xinyuan Hu, Yutang Lin, Fengli XuWWW 2025 · 31 citations
- WebVoyager: Building an End-to-End Web Agent with Large Multimodal ModelsHongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu et al.ACL 2024 · 30 citations
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web TasksJing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur et al.ACL 2024 · 25 citations
Related papers
- Next-Gen CAPTCHAs: Leveraging the Cognitive Gap for Scalable and Diverse GUI-Agent DefenseJiacheng Liu, Yaxin Luo, Jiacheng Cui, Xinyi Shang et al.ICML 2026 · 2 citations
- GTA: Generating Long-horizon Tasks for Web Agents at ScaleTenghao Huang, Kung-Hsiang Huang, Prafulla Kumar Choubey, Yilun Zhou et al.ACL 2026
- AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World ContextsKeyu Li, Junhao Shi, Yang Xiao, Mohan Jiang et al.ACL 2026 · 14 citations
- CAPTCHA Solving for Native GUI Agents: Automated Reasoning-Action Data Generation and Self-Corrective TrainingYuxi Chen, Haoyu Zhai, Chenkai Wang, Rui Yang et al.ICML 2026
- AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?Ori Yoran, Samuel Joseph Amouyal, Chaitanya Malaviya, Ben Bogin et al.EMNLP 2024 · 5 citations
