Closing the Spatial Execution Gap in Digital Whiteboards via Verifiable Reinforcement Learning
Chang Liu, Benjamin Wagley, Zibo Wang, Mehmet E. Belviranli, Bo Wu
Abstract
While multi-modal large language models such as GPT-5 demonstrate exceptional general understanding, they suffer from a fundamental Spatial Execution Gap, failing to translate visual semantics into precise, schema-valid coordinate operations in interactive environments. In this work, we show that model scale alone cannot close this gap; instead, verifiable structured reasoning provides the key to spatial precision. We present a comprehensive pipeline that leverages Group Relative Policy Optimization to enforce a strict Identify-Reason-Verify protocol, effectively shifting the computational burden from parameters to test-time reasoning. By utilizing a multi-agent system to distill optimal reasoning schemas and training on execution-verifiable rewards, our specialized 3B agent achieves 100% format coherence and 81.12% operation accuracy on digital whiteboard tasks. Crucially, our approach outperforms a state-of-the-art frontier model, GPT-5, by 16.75% in operation accuracy. The results suggest that for complex user interface manipulation, small, RL-aligned models with dedicated reasoning protocols are superior to generalist frontier models, offering a promising direction for building reliable web agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d6b0b1b-495f-4581-8fb4-2e5de44c0322Builds on15
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye et al.ICLR 2026 · 670 citations
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo et al.NeurIPS 2025 · 528 citations
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo et al.NeurIPS 2024 · 412 citations
- Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language ModelsJiayu Wang, Yifei Ming, Zhenmei Shi, Vibhav Vineet et al.NeurIPS 2024 · 166 citations
Related papers
- 3D-RFT: Reinforcement Fine-Tuning for Video-based 3D Scene UnderstandingXiongkun Linghu, Jiangyong Huang, Baoxiong Jia, Siyuan HuangICML 2026 · 1 citation
- CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy OptimizationXinhai Hou, Shaoyuan Xu, Manan Biyani, Moyan Li et al.CVPR 2026 · 25 citations
- Geometrically-Constrained Agent for Spatial ReasoningZeren Chen, Xiaoya Lu, Zhijie Zheng, Pengrui Li et al.CVPR 2026 · 29 citations
- The Art of Interrogation: Consistency Amplifies Factuality in Spatial ReasoningThéo Uscidda, Marta Gazulla, Maks Ovsjanikov, Federico Tombari et al.ICML 2026
- UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement LearningZhengxi Lu, Yuxiang Chai, Yaxuan Guo, Xi Yin et al.AAAI 2026 · 103 citations
