Learning to Edit Visual Programs with Self-Supervision
R. Kenny Jones, Renhao Zhang, Aditya Ganeshan, Daniel Ritchie
Abstract
We design a system that learns how to edit visual programs. Our edit network consumes a complete input program and a visual target. From this input, we task our network with predicting a local edit operation that could be applied to the input program to improve its similarity to the target. In order to apply this scheme for domains that lack program annotations, we develop a self-supervised learning approach that integrates this edit network into a bootstrapped finetuning loop along with a network that predicts entire programs in one-shot. Our joint finetuning scheme, when coupled with an inference procedure that initializes a population from the one-shot model and evolves members of this population with the edit network, helps to infer more accurate visual programs. Over multiple domains, we experimentally compare our method against the alternative of using only the one-shot model, and find that even under equal search-time budgets, our editing-based paradigm provides significant advantages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c8cbeeb0-aa55-4dc0-afff-a35953dc3d84Cited by top-tier papers2
- Residual Primitive Fitting of 3D Shapes with SuperFrustaAditya Ganeshan, Matheus Gadelha, Thibault Groueix, Zhiqin Chen et al.CVPR 2026 · 5 citations
- Pattern Analogies: Learning to Perform Programmatic Image Edits by AnalogyAditya Ganeshan, Thibault Groueix, Paul Guerrero, Radomír Mech et al.CVPR 2025
Builds on25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- DeepCAD: A Deep Generative Network for Computer-Aided Design ModelsRundi Wu, Chang Xiao, Changxi ZhengICCV 2021 · 290 citations
Related papers
- Improving Unsupervised Visual Program Inference with Code Rewriting FamiliesAditya Ganeshan, R. Kenny Jones, Daniel RitchieICCV 2023 · 13 citations
- Self-Training Large Language Models for Improved Visual Program Synthesis With Visual ReinforcementZaid Khan, Vijay Kumar B. G, Samuel Schulter, Yun Fu et al.CVPR 2024
- Program-Guided Image ManipulatorsXiuming Zhang, Jiayuan Mao, Yikai Li, William T. Freeman et al.ICCV 2019 · 25 citations
- Visual Programming: Compositional visual reasoning without trainingTanmay Gupta, Aniruddha KembhaviCVPR 2023
- Learning Where to Edit Vision TransformersYunqiao Yang, Long-Kai Huang, Shengzhuang Chen, Kede Ma et al.NeurIPS 2024 · 6 citations
