GeoLoom: High-quality Geometric Diagram Generation from Textual Input
Xiaojing Wei, Ting Zhang, Wei He, Jingdong Wang, Hua Huang
Abstract
High-quality geometric diagram generation presents both a challenge and an opportunity: it demands strict spatial accuracy while offering well-defined constraints to guide generation. Inspired by recent advances in geometry problem solving that employ formal languages and symbolic solvers for enhanced correctness and interpretability, we propose GeoLoom, a novel framework for text-to-diagram generation in geometric domains. GeoLoom comprises two core components: an autoformalization module that translates natural language into a specifically designed generation-oriented formal language Ge-oLingua, and a coordinate solver that maps formal constraints to precise coordinates using the efficient Monte Carlo optimization. To support this framework, we introduce GeoNF, a dataset aligning natural language geometric descriptions with formal GeoLingua descriptions. We further propose a constraint-based evaluation metric that quantifies structural deviation, offering mathematically grounded supervision for iterative refinement. Empirical results demonstrate that Ge-oLoom significantly outperforms state-of-the-art baselines in structural fidelity, providing a principled foundation for interpretable and scalable diagram generation. The dataset is publicly available at: github.com/BNU-ERC-ITEA/GeoNF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 081be346-855c-4b01-982b-a85d2c9b5bb0Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- DeepSVG: A Hierarchical Generative Network for Vector Graphics AnimationAlexandre Carlier, Martin Danelljan, Alexandre Alahi, Radu TimofteNeurIPS 2020 · 247 citations
- Emu: Generative Pretraining in MultimodalityQuan Sun, Qiying Yu, Yufeng Cui, Fan Zhang et al.ICLR 2024 · 161 citations
Related papers
- Enhancing the Geometric Problem-Solving Ability of Multimodal LLMs via Symbolic-Neural IntegrationYicheng Pan, Zhenrong Zhang, Pengfei Hu, Jiefeng Ma et al.ACM MM 2025 · 3 citations
- Synthesizing Multimodal Geometry Datasets from Scratch and Enabling Visual Alignment via Plotting CodeHaobo Lin, Tianyi Bai, Chen Chen, Jiajun Zhang et al.ICML 2026 · 1 citation
- Autoformalizing Euclidean GeometryLogan Murphy, Kaiyu Yang, Jialiang Sun, Zhaoyu Li et al.ICML 2024 · 16 citations
- Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic ReasoningPan Lu, Ran Gong, Shibiao Jiang, Liang Qiu et al.ACL 2021
- Euclean: Automated Geometry Problem Formalization with Unified Verification in LeanLinbin Tang, Jingyan You, Zilin Kang, Hanzhang Liu et al.ICML 2026 · 1 citation
