GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
Renqiu Xia, Mingsheng Li, Hancheng Ye, Wenjie Wu, Hongbin Zhou, Jiakang Yuan, Tianshuo Peng, Xinyu Cai, Xiangchao Yan, Bin Wang, Conghui He, Botian Shi
摘要
Despite their proficiency in general tasks, Multi-modal Large Language Models (MLLMs) struggle with automatic Geometry Problem Solving (GPS), which demands understanding diagrams, interpreting symbols, and performing complex reasoning. This limitation arises from their pre-training on natural images and texts, along with the lack of automated verification in the problem-solving process. Besides, current geometric specialists are limited by their task-specific designs, making them less effective for broader geometric problems. To this end, we present GeoX, a multi-modal large model focusing on geometric understanding and reasoning tasks. Given the significant differences between geometric diagram-symbol and natural image-text, we introduce unimodal pre-training to develop a diagram encoder and symbol decoder, enhancing the understanding of geometric images and corpora. Furthermore, we introduce geometry-language alignment, an effective pre-training paradigm that bridges the modality gap between unimodal geometric experts. We propose a Generator-And-Sampler Transformer (GS-Former) to generate discriminative queries and eliminate uninformative representations from unevenly distributed geometric signals. Finally, GeoX benefits from visual instruction tuning, empowering it to take geometric images and questions as input and generate verifiable solutions. Experiments show that GeoX outperforms both generalists and geometric specialists on publicly recognized benchmarks, such as GeoQA, UniGeo, Geometry3K, and PGPS9k. Our code is available at https://github.com/Alpha-Innovator/GeoX Published as a conference paper at ICLR 2025 Two Examples Comparing our GeoX and GPT-4V Two Examples Comparing our GeoX and GPT-4V If angle D is 50°, since the line from O to D would create an isosceles triangle with OC and OD being equal radii of the circle, angle ∠COD would correspond to the central angle subtending arc CD. Therefore, inscribed angle ∠CAD would be half the measure of the central angle ∠COD because an inscribed angle is half the measure of the corresponding central angle. Given ∠D is 50°, this means ∠CAD would be 25°. So, angle A (∠CAD) is 25.0°.The correct answer from the choices given is 20.0. g_minus C_3 C_2, g_minus V_0 N_0, g_half V_1 20
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Unlocking Multimodal Mathematical Reasoning via Process Reward ModelRuilin Luo, Zhuofan Zheng, Lei Wang, Yifan Wang 等NeurIPS 2025 · 被引用 38 次
- PixelCraft: A Multi-Agent system for High-Fidelity Visual Reasoning on Structured ImagesShuoshuo Zhang, Zijian Li, Yizhen Zhang, Jingjing Fu 等ICLR 2026 · 被引用 12 次
- Empowering Small VLMs to Think with Dynamic Memorization and ExplorationJiazhen Liu, Yuchuan Deng, Long ChenICLR 2026 · 被引用 7 次
- A Survey of Deep Learning for Geometry Problem SolvingJianzhe Ma, Wenxuan Wang, Qin JinACL 2026 · 被引用 5 次
- GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical EvaluationYuan Feng, Yue Yang, Xiaohan He, Jiatong Zhao 等ICLR 2026 · 被引用 4 次
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- Enhancing the Geometric Problem-Solving Ability of Multimodal LLMs via Symbolic-Neural IntegrationYicheng Pan, Zhenrong Zhang, Pengfei Hu, Jiefeng Ma 等ACM MM 2025 · 被引用 3 次
- GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary LinesYumeng Fu, Jiayin Zhu, Lingling Zhang, Wenjun Wu 等ACL 2026 · 被引用 3 次
- Synthesizing Multimodal Geometry Datasets from Scratch and Enabling Visual Alignment via Plotting CodeHaobo Lin, Tianyi Bai, Chen Chen, Jiajun Zhang 等ICML 2026 · 被引用 1 次
- GNS: Solving Plane Geometry Problems by Neural-Symbolic Reasoning with Multi-Modal LLMsMaizhen Ning, Zihao Zhou, Qiufeng Wang, Xiaowei Huang 等AAAI 2025 · 被引用 10 次
- GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal ModelsJingxuan Wei, Caijun Jia, Xi Bai, Xinglong Xu 等CVPR 2026 · 被引用 7 次
