Enhancing Geometric Perception in VLMs via Translator-Guided Reinforcement Learning
Hao Yu, Shuning Jia, Guanghao Li, Wenhao Jiang, Chun Yuan
摘要
Vision-language models (VLMs) often struggle with geometric reasoning due to their limited perception of fundamental diagram elements. To tackle this challenge, we introduce GeoPerceive, a benchmark comprising diagram instances paired with domain-specific language (DSL) representations, along with an efficient automatic data generation pipeline. This design enables the isolated evaluation of geometric perception independently from reasoning. To exploit the data provided by GeoPerceive for enhancing the geometric perception capabilities of VLMs, we propose GeoDPO, a translator-guided reinforcement learning framework. GeoDPO employs an NL-to-DSL translator, which is trained on synthetic pairs generated by the data engine of GeoPerceive, to bridge natural language and DSL. This translator facilitates the computation of fine-grained, DSL-level scores, which serve as reward signals in reinforcement learning. We assess GeoDPO on both in-domain and out-of-domain datasets, spanning tasks in geometric perception as well as downstream reasoning. Experimental results demonstrate that, while supervised fine-tuning (SFT) offers only marginal improvements and may even impair performance in out-of-domain scenarios, GeoDPO achieves substantial gains: on in-domain data, on out-of-domain data, and on downstream reasoning tasks. These findings underscore the superior performance and generalization ability of GeoDPO over SFT. All codes are released at https://github.com/Longin-Yu/GeoPerceive to ensure reproducibility.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation GenerationZihan Su, Hongyang Wei, Kangrui Cen, Yong Wang 等ICML 2026 · 被引用 15 次
- PromptHub: Enhancing Multi-Prompt Visual In-Context Learning with Locality-Aware Fusion, Concentration and AlignmentTianci Luo, Jinpeng Wang, Shiyu Qin, Niu Lian 等ICLR 2026 · 被引用 3 次
它引用的顶会 Paper8
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical ReasoningPan Lu, Liang Qiu, Kai-Wei Chang, Ying Nian Wu 等ICLR 2023 · 被引用 41 次
- UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical ExpressionJiaqi Chen, Tong Li, Jinghui Qin, Pan Lu 等EMNLP 2022 · 被引用 37 次
- A Symbolic Characters Aware Model for Solving Geometry ProblemsMaizhen Ning, Qiu-Feng Wang, Kaizhu Huang, Xiaowei HuangACM MM 2023 · 被引用 18 次
相关 Paper
- Synthesizing Multimodal Geometry Datasets from Scratch and Enabling Visual Alignment via Plotting CodeHaobo Lin, Tianyi Bai, Chen Chen, Jiajun Zhang 等ICML 2026 · 被引用 1 次
- GeoTikzBridge: Advancing Multimodal Code Generation for Geometric Perception and ReasoningJiayin Sun, Caixia Sun, Boyu Yang, hailin li 等CVPR 2026 · 被引用 3 次
- GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical EvaluationYuan Feng, Yue Yang, Xiaohan He, Jiatong Zhao 等ICLR 2026 · 被引用 4 次
- Integrating Visual Interpretation and Linguistic Reasoning for Geometric Problem SolvingZixian Guo, Ming Liu, Qilong Wang, Zhilong Ji 等ICCV 2025 · 被引用 1 次
- CReFT-CAD: Boosting Orthographic Projection Reasoning for CAD via Reinforcement Fine-TuningKe Niu, Zhuofan Chen, Haiyang Yu, Yuwen Chen 等NeurIPS 2025 · 被引用 8 次
