Lune

ACL2026顶会

MirrorCAPTCHA: Wild CAPTCHA, Wild Distribution, Wild Web-based Platform Meet Multimodal LLM Agents

Xiangyu Wu, Yuwei Hu, Tianyu Cui, Yueying Tian, Qingguo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, Yang Yang, Jianfeng Lu

2026年份

摘要

The path to fully autonomous web agents is currently hindered by a critical bottleneck: their limited ability to handle CAPTCHA. Existing agent benchmarks largely ignore this practical challenge, failing to evaluate an agent’s real-world capacity to solve CAPTCHA. To bridge this gap, we conduct a comprehensive analysis of real-world CAPTCHA distributions and introduce MirrorCAPTCHA , a benchmark annotated with Weighted Pass Rate and a newly proposed metric Completion Degree . Mirror-CAPTCHA is designed to serve as a “mirror” that faithfully reflects the automation capabilities of agents in real scenarios. We filter 2 , 095 websites from Common Crawl, identify the CAPTCHA deployed on these sites, and cluster them into 18 distinct categories us-ing K-means algorithm. To ensure practicality, we extract a web subgraph from Common Crawl covering these websites and use random walks to simulate real-world CAPTCHA encounter frequencies, yielding a realistic measure of agents’ ability. Additionally, we develop a lightweight synthetic data pipeline to train Ovis2-Agent-CAPTCHA-8B , which significantly outperforms current state-of-the-art closed-source models on MirrorCAPTCHA, achieving a 9 . 4% higher average Weighted Pass Rate and a 2 . 13% higher average Completion Degree than the runner-up, Gemini-2.5-Pro .

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper7

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖