Tikzero: Zero-Shot Text-Guided Graphics Program Synthesis
Jonas Belouadi, Eddy Ilg, Margret Keuper, Hideki Tanaka, Masao Utiyama, Raj Dabre, Steffen Eger, Simone Paolo Ponzetto
Abstract
Automatically synthesizing figures from text captions is a compelling capability. However, achieving high geometric precision and editability requires representing figures as graphics programs in languages like TikZ, and aligned training data (i.e., graphics programs with captions) remains scarce. Meanwhile, large amounts of unaligned graphics programs and captioned raster images are more readily available. We reconcile these disparate data sources by presenting TikZero, which decouples graphics program generation from text understanding by using image representations as an intermediary bridge. It enables independent training on graphics programs and captioned images and allows for zero-shot text-guided graphics program synthesis during inference. We show that our method substantially outperforms baselines that can only operate with caption-aligned graphics programs. Furthermore, when leveraging caption-aligned graphics programs as a complementary training signal, or exceeds the performance of much larger models, including commercial systems like GPT-4o. Our code, datasets, and select models are publicly available.11https://github.com/potamides/DeTikZify
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Rendering-Aware Reinforcement Learning for Vector Graphics GenerationJuan A. Rodríguez, Haotian Zhang, Abhay Puri, Rishav Pramanik et al.NeurIPS 2025 · 42 citations
- AutoFigure: Generating and Refining Publication-Ready Scientific IllustrationsMinjun Zhu, Zhen Lin, Yixuan Weng, Panzhong Lu et al.ICLR 2026 · 28 citations
- TikZilla: Scaling Text-to-TikZ with High-Quality Data and Reinforcement LearningChristian Greisinger, Steffen EgerICLR 2026 · 5 citations
- GeoTikzBridge: Advancing Multimodal Code Generation for Geometric Perception and ReasoningJiayin Sun, Caixia Sun, Boyu Yang, hailin li et al.CVPR 2026 · 3 citations
- Dynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From VideosChia-Hsiang Kao, Cong Phuoc Huynh, Chien-Yi Wang, Noranart Vesdapunt et al.CVPR 2026 · 2 citations
Builds on38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- DeTikZify: Synthesizing Graphics Programs for Scientific Figures and Sketches with TikZJonas Belouadi, Simone Paolo Ponzetto, Steffen EgerNeurIPS 2024
- AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZJonas Belouadi, Anne Lauscher, Steffen EgerICLR 2024 · 64 citations
- Altogether: Image Captioning via Re-aligning Alt-textHu Xu, Po-Yao Huang, Xiaoqing Ellen Tan, Ching-Feng Yeh et al.EMNLP 2024 · 3 citations
- Variational Distribution Learning for Unsupervised Text-to-Image GenerationMinsoo Kang, Doyup Lee, Jiseob Kim, Saehoon Kim et al.CVPR 2023
- CLIP-Forge: Towards Zero-Shot Text-to-Shape GenerationAditya Sanghi, Hang Chu, Joseph G. Lambourne, Ye Wang et al.CVPR 2022 · 206 citations
