Diffusion On Syntax Trees For Program Synthesis
Shreyas Kapur, Erik Jenner, Stuart Russell
Abstract
Large language models generate code one token at a time. Their autoregressive generation process lacks the feedback of observing the program's output. Training LLMs to suggest edits directly can be challenging due to the scarcity of rich edit data. To address these problems, we propose neural diffusion models that operate on syntax trees of any context-free grammar. Similar to image diffusion models, our method also inverts "noise" applied to syntax trees. Rather than generating code sequentially, we iteratively edit it while preserving syntactic validity, which makes it easy to combine this neural model with search. We apply our approach to inverse graphics tasks, where our model learns to convert images into programs that produce those images. Combined with search, our model is able to write graphics programs, see the execution result, and debug them to meet the required specifications. We additionally show how our system can write graphics programs for hand-drawn sketches. Video results can be found at https://td-anon.github.io . Introduction Large language models (LLMs) have made remarkable progress in code generation, but their autoregressive nature presents a fundamental challenge: they generate code token by token, without access to the program's runtime output from the previously generated tokens. This makes it difficult to correct errors, as the model lacks the feedback loop of seeing the program's output and adjusting accordingly. While LLMs can be trained to suggest edits to existing code [6, 42, 17] , acquiring sufficient training data for this task is difficult. In this paper, we introduce a new approach to program synthesis using neural diffusion models that operate directly on syntax trees. Diffusion models have previously been used to great success in image generation [14, 22, 31] . By leveraging diffusion, we let the model learn to iteratively refine programs while ensuring syntactic validity. Crucially, our approach allows the model to observe the program's output at each step, effectively enabling a debugging process. In the spirit of systems like AlphaZero [29], the iterative nature of diffusion naturally lends itself to search-based program synthesis. By training a value model alongside our diffusion model, we can guide the denoising process toward programs that are likely to achieve the desired output. This allows us to efficiently explore the program space, making more informed decisions at each step of the generation process. We implement our approach for inverse graphics tasks, where we posit domain-specific languages for drawing images. Inverse graphics tasks are naturally suitable for our approach since small changes in the code produce semantically meaningful changes in the rendered image. For example, a misplaced shape on the image can be easily seen and fixed in program space.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang et al.NeurIPS 2025 · 949 citations
- Tikzero: Zero-Shot Text-Guided Graphics Program SynthesisJonas Belouadi, Eddy Ilg, Margret Keuper, Hideki Tanaka et al.ICCV 2025 · 24 citations
- MultiMat: Multimodal Program Synthesis for Procedural Materials using Large Multimodal ModelsJonas Belouadi, Tamy Boubekeur, Adrien KaiserICLR 2026 · 2 citations
Builds on12
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- Multimodal Few-Shot Learning with Frozen Language ModelsMaria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami et al.NeurIPS 2021 · 1,020 citations
- Equivariant Diffusion for Molecule Generation in 3DEmiel Hoogeboom, Victor Garcia Satorras, Clément Vignac, Max WellingICML 2022 · 865 citations
Related papers
- Constrained Decoding of Diffusion LLMs with Context-Free GrammarsNiels Mündler, Jasper Dekoninck, Martin VechevICLR 2026 · 18 citations
- D-AR: Diffusion via Autoregressive ModelsZiteng Gao, Mike Zheng ShouICLR 2026 · 11 citations
- RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement LearningJonas Gehring, Kunhao Zheng, Jade Copet, Vegard Mella et al.ICML 2025
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang et al.NeurIPS 2022 · 1,546 citations
- Lookahead-Then-Verify: Reliable Constrained Decoding for Diffusion LLMs under Context-Free GrammarsYitong Zhang, Yongmin Li, Yuetong Liu, Jia Li et al.ISSTA 2026
