GeoGPT4V: Towards Geometric Multi-modal Large Language Models with Geometric Image Generation
Shihao Cai, Keqin Bao, Hangyu Guo, Jizhi Zhang, Jun Song, Bo Zheng
摘要
Large language models have seen widespread adoption in math problem-solving. However, in geometry problems that usually require visual aids for better understanding, even the most advanced multi-modal models currently still face challenges in effectively using image information. High-quality data is crucial for enhancing the geometric capabilities of multi-modal models, yet existing open-source datasets and related efforts are either too challenging for direct model learning or suffer from misalignment between text and images. To overcome this issue, we introduce a novel pipeline that leverages GPT-4 and GPT-4V to generate relatively basic geometry problems with aligned text and images, facilitating model learning. We have produced a dataset of 4.9K geometry problems and combined it with 19K opensource data to form our GeoGPT4V dataset. Experimental results demonstrate that the Ge-oGPT4V dataset significantly improves the geometry performance of various models on the MathVista and MathVision benchmarks. The code is available at https://github.com/ Lanyu0303/GeoGPT4V_Project .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical ReasoningWenwen Zhuang, Xin Huang, Xiantao Zhang, Jin ZengAAAI 2025 · 被引用 66 次
- Unlocking Multimodal Mathematical Reasoning via Process Reward ModelRuilin Luo, Zhuofan Zheng, Lei Wang, Yifan Wang 等NeurIPS 2025 · 被引用 38 次
- A Survey of Deep Learning for Geometry Problem SolvingJianzhe Ma, Wenxuan Wang, Qin JinACL 2026 · 被引用 5 次
- Enhancing the Geometric Problem-Solving Ability of Multimodal LLMs via Symbolic-Neural IntegrationYicheng Pan, Zhenrong Zhang, Pengfei Hu, Jiefeng Ma 等ACM MM 2025 · 被引用 3 次
- GeoUni: A Unified Model for Generating Geometry Diagrams, Problems and Problem SolutionsJo-Ku Cheng, Zeren Zhang, Ran Chen, Jingyang Deng 等ACM MM 2025 · 被引用 2 次
它引用的顶会 Paper18
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer 等NeurIPS 2023 · 被引用 1,486 次
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras 等EMNLP 2021 · 被引用 937 次
相关 Paper
- G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language ModelJiahui Gao, Renjie Pi, Jipeng Zhang, Jiacheng Ye 等ICLR 2025 · 被引用 5 次
- GNS: Solving Plane Geometry Problems by Neural-Symbolic Reasoning with Multi-Modal LLMsMaizhen Ning, Zihao Zhou, Qiufeng Wang, Xiaowei Huang 等AAAI 2025 · 被引用 10 次
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu 等ICLR 2024 · 被引用 1,472 次
- It Ain't Over: A Multi-aspect Diverse Math Word Problem DatasetJiwoo Kim, Youngbin Kim, Ilwoong Baek, JinYeong Bak 等EMNLP 2023 · 被引用 2 次
- Primitive Vision: Improving Diagram Understanding in MLLMsShan Zhang, Aotian Chen, Yanpeng Sun, Jindong Gu 等ICML 2025
