Spatial CAPTCHA: Generatively Benchmarking Spatial Reasoning for Human-Machine Differentiation
Arina Kharlamova, Bowei He, Chen Ma, Xue Liu
Abstract
Online services rely on CAPTCHAs as a first line of defense against automated abuse, yet recent advances in multi-modal large language models (MLLMs) have eroded the effectiveness of conventional designs that focus on text recognition or 2D image understanding. To address this challenge, we present Spatial CAPTCHA, a novel human-verification framework that leverages fundamental differences in spatial reasoning between humans and MLLMs. Unlike existing CAPTCHAs which rely on low-level perception tasks that are vulnerable to modern AI, Spatial CAPTCHA generates dynamic questions requiring geometric reasoning, perspective-taking, occlusion handling, and mental rotation. These skills are intuitive for humans but difficult for state-of-the-art (SOTA) AI systems. The system employs a procedural generation pipeline with constraint-based difficulty control, automated correctness verification, and human-in-the-loop validation to ensure scalability, robustness, and adaptability. Evaluation on a corresponding benchmark, Spatial-CAPTCHA-Bench, demonstrates that humans vastly outperform 10 state-of-the-art MLLMs, with the best model achieving only 31.0% Pass@1 accuracy. Furthermore, we compare Spatial CAPTCHA with Google reCAPTCHA, which confirms its effectiveness as both a security mechanism and a diagnostic tool for spatial reasoning in AI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08201374-aea8-4ee7-aaee-583dfbd4b9b8Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial IntelligenceDiankun Wu, Fangfu Liu, Yi-Hsin Hung, Yueqi DuanNeurIPS 2025 · 245 citations
Related papers
- IllusionCAPTCHA: A CAPTCHA based on Visual IllusionZiqi Ding, Gelei Deng, Yi Liu, Junchen Ding et al.WWW 2025 · 15 citations
- Oedipus: LLM-enchanced Reasoning CAPTCHA SolverGelei Deng, Haoran Ou, Yi Liu, Jie Zhang et al.CCS 2025
- COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA SolversJunyu Wang, Changjia Zhu, Yuanbo Zhou, Lingyao Li et al.USENIX Security 2026 · 4 citations
- MMSI-Bench: A Benchmark for Multi-Image Spatial IntelligenceSihan Yang, Runsen Xu, Yiman Xie, Sizhe Yang et al.ICLR 2026 · 195 citations
- Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual SimulationsLinjie Li, Mahtab Bigverdi, Jiawei Gu, Zixian Ma et al.ICLR 2026 · 22 citations
