GeoGPT4V: Towards Geometric Multi-modal Large Language Models with Geometric Image Generation
Shihao Cai, Keqin Bao, Hangyu Guo, Jizhi Zhang, Jun Song, Bo Zheng
Abstract
Large language models have seen widespread adoption in math problem-solving. However, in geometry problems that usually require visual aids for better understanding, even the most advanced multi-modal models currently still face challenges in effectively using image information. High-quality data is crucial for enhancing the geometric capabilities of multi-modal models, yet existing open-source datasets and related efforts are either too challenging for direct model learning or suffer from misalignment between text and images. To overcome this issue, we introduce a novel pipeline that leverages GPT-4 and GPT-4V to generate relatively basic geometry problems with aligned text and images, facilitating model learning. We have produced a dataset of 4.9K geometry problems and combined it with 19K opensource data to form our GeoGPT4V dataset. Experimental results demonstrate that the Ge-oGPT4V dataset significantly improves the geometry performance of various models on the MathVista and MathVision benchmarks. The code is available at https://github.com/ Lanyu0303/GeoGPT4V_Project .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd23a512-e762-47d3-a998-f06d299dcd65Cited by top-tier papers12
- Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical ReasoningWenwen Zhuang, Xin Huang, Xiantao Zhang, Jin ZengAAAI 2025 · 66 citations
- Unlocking Multimodal Mathematical Reasoning via Process Reward ModelRuilin Luo, Zhuofan Zheng, Lei Wang, Yifan Wang et al.NeurIPS 2025 · 38 citations
- A Survey of Deep Learning for Geometry Problem SolvingJianzhe Ma, Wenxuan Wang, Qin JinACL 2026 · 5 citations
- Enhancing the Geometric Problem-Solving Ability of Multimodal LLMs via Symbolic-Neural IntegrationYicheng Pan, Zhenrong Zhang, Pengfei Hu, Jiefeng Ma et al.ACM MM 2025 · 3 citations
- GeoUni: A Unified Model for Generating Geometry Diagrams, Problems and Problem SolutionsJo-Ku Cheng, Zeren Zhang, Ran Chen, Jingyang Deng et al.ACM MM 2025 · 2 citations
Builds on18
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer et al.NeurIPS 2023 · 1,486 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
Related papers
- G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language ModelJiahui Gao, Renjie Pi, Jipeng Zhang, Jiacheng Ye et al.ICLR 2025 · 5 citations
- GNS: Solving Plane Geometry Problems by Neural-Symbolic Reasoning with Multi-Modal LLMsMaizhen Ning, Zihao Zhou, Qiufeng Wang, Xiaowei Huang et al.AAAI 2025 · 10 citations
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
- It Ain't Over: A Multi-aspect Diverse Math Word Problem DatasetJiwoo Kim, Youngbin Kim, Ilwoong Baek, JinYeong Bak et al.EMNLP 2023 · 2 citations
- Primitive Vision: Improving Diagram Understanding in MLLMsShan Zhang, Aotian Chen, Yanpeng Sun, Jindong Gu et al.ICML 2025
