GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
Renqiu Xia, Mingsheng Li, Hancheng Ye, Wenjie Wu, Hongbin Zhou, Jiakang Yuan, Tianshuo Peng, Xinyu Cai, Xiangchao Yan, Bin Wang, Conghui He, Botian Shi
Abstract
Despite their proficiency in general tasks, Multi-modal Large Language Models (MLLMs) struggle with automatic Geometry Problem Solving (GPS), which demands understanding diagrams, interpreting symbols, and performing complex reasoning. This limitation arises from their pre-training on natural images and texts, along with the lack of automated verification in the problem-solving process. Besides, current geometric specialists are limited by their task-specific designs, making them less effective for broader geometric problems. To this end, we present GeoX, a multi-modal large model focusing on geometric understanding and reasoning tasks. Given the significant differences between geometric diagram-symbol and natural image-text, we introduce unimodal pre-training to develop a diagram encoder and symbol decoder, enhancing the understanding of geometric images and corpora. Furthermore, we introduce geometry-language alignment, an effective pre-training paradigm that bridges the modality gap between unimodal geometric experts. We propose a Generator-And-Sampler Transformer (GS-Former) to generate discriminative queries and eliminate uninformative representations from unevenly distributed geometric signals. Finally, GeoX benefits from visual instruction tuning, empowering it to take geometric images and questions as input and generate verifiable solutions. Experiments show that GeoX outperforms both generalists and geometric specialists on publicly recognized benchmarks, such as GeoQA, UniGeo, Geometry3K, and PGPS9k. Our code is available at https://github.com/Alpha-Innovator/GeoX Published as a conference paper at ICLR 2025 Two Examples Comparing our GeoX and GPT-4V Two Examples Comparing our GeoX and GPT-4V If angle D is 50°, since the line from O to D would create an isosceles triangle with OC and OD being equal radii of the circle, angle ∠COD would correspond to the central angle subtending arc CD. Therefore, inscribed angle ∠CAD would be half the measure of the central angle ∠COD because an inscribed angle is half the measure of the corresponding central angle. Given ∠D is 50°, this means ∠CAD would be 25°. So, angle A (∠CAD) is 25.0°.The correct answer from the choices given is 20.0. g_minus C_3 C_2, g_minus V_0 N_0, g_half V_1 20
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- Unlocking Multimodal Mathematical Reasoning via Process Reward ModelRuilin Luo, Zhuofan Zheng, Lei Wang, Yifan Wang et al.NeurIPS 2025 · 38 citations
- PixelCraft: A Multi-Agent system for High-Fidelity Visual Reasoning on Structured ImagesShuoshuo Zhang, Zijian Li, Yizhen Zhang, Jingjing Fu et al.ICLR 2026 · 12 citations
- Empowering Small VLMs to Think with Dynamic Memorization and ExplorationJiazhen Liu, Yuchuan Deng, Long ChenICLR 2026 · 7 citations
- A Survey of Deep Learning for Geometry Problem SolvingJianzhe Ma, Wenxuan Wang, Qin JinACL 2026 · 5 citations
- GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical EvaluationYuan Feng, Yue Yang, Xiaohan He, Jiatong Zhao et al.ICLR 2026 · 4 citations
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Enhancing the Geometric Problem-Solving Ability of Multimodal LLMs via Symbolic-Neural IntegrationYicheng Pan, Zhenrong Zhang, Pengfei Hu, Jiefeng Ma et al.ACM MM 2025 · 3 citations
- GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary LinesYumeng Fu, Jiayin Zhu, Lingling Zhang, Wenjun Wu et al.ACL 2026 · 3 citations
- Synthesizing Multimodal Geometry Datasets from Scratch and Enabling Visual Alignment via Plotting CodeHaobo Lin, Tianyi Bai, Chen Chen, Jiajun Zhang et al.ICML 2026 · 1 citation
- GNS: Solving Plane Geometry Problems by Neural-Symbolic Reasoning with Multi-Modal LLMsMaizhen Ning, Zihao Zhou, Qiufeng Wang, Xiaowei Huang et al.AAAI 2025 · 10 citations
- GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal ModelsJingxuan Wei, Caijun Jia, Xi Bai, Xinglong Xu et al.CVPR 2026 · 7 citations
