Block and Detail: Scaffolding Sketch-to-Image Generation
Vishnu Sarukkai, Lu Yuan, Mia Tang, Maneesh Agrawala, Kayvon Fatahalian
Abstract
We introduce a novel sketch-to-image tool that aligns with the iterative refinement process of artists. Our tool lets users sketch blocking strokes to coarsely represent the placement and form of objects and detail strokes to refine their shape and silhouettes. We develop a two-pass algorithm for generating high-fidelity images from such sketches at any point in the iterative process. In the first pass we use a ControlNet to generate an image that strictly follows all the strokes (blocking and detail) and in the second pass we add variation by renoising regions surrounding blocking strokes. We also present a dataset generation scheme that, when used to train a ControlNet architecture, allows regions that do not contain strokes to be interpreted as not-yet-specified regions rather than empty space. We show that this partial-sketch-aware ControlNet can generate coherent elements from partial sketches that only contain a small number of strokes. The high-fidelity images produced by our approach serve as scaffolds that can help the user adjust the shape and proportions of objects or add additional elements to the composition. We demonstrate the effectiveness of our approach with a variety of examples and evaluative comparisons. Quantitatively, evaluative user feedback indicates that novice viewers prefer the quality of images from our algorithm over a baseline Scribble ControlNet for 84% of the pairs and found our images had less distortion in 81% of the pairs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e235845-92d0-4370-84d8-44836d9e0f3bCited by top-tier papers7
- DesignWeaver: Dimensional Scaffolding for Text-to-Image Product DesignSirui Tao, Ivan Liang, Cindy Peng, Zhiqing Wang et al.CHI 2025 · 19 citations
- Brickify: Enabling Expressive Design Intent Specification through Direct Manipulation on Design TokensXinyu Shi, Yinghou Wang, Ryan A. Rossi, Jian ZhaoCHI 2025 · 18 citations
- GenPara: Enhancing the 3D Design Editing Process by Inferring Users' Regions of Interest with Text-Conditional Shape ParametersJiin Choi, Seung Won Lee, Kyung Hoon HyunCHI 2025 · 6 citations
- DesignTrace: Exploring, Iterating and Tracking Design Alternatives with GenAIXiaohan Peng, Debanjana Haldar, Wendy E. Mackay, Janin KochCHI 2026 · 3 citations
- Protosampling: Enabling Free-Form Convergence of Sampling and Prototyping through Canvas-Driven Visual AI GenerationAlicia Guo, David Ledo, George W. Fitzmaurice, Fraser AndersonCHI 2026 · 2 citations
Builds on10
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsChenlin Meng, Yutong He, Yang Song, Jiaming Song et al.ICLR 2022 · 2,128 citations
- MultiDiffusion: Fusing Diffusion Paths for Controlled Image GenerationOmer Bar-Tal, Lior Yariv, Yaron Lipman, Tali DekelICML 2023 · 575 citations
- Blended Latent DiffusionOmri Avrahami, Ohad Fried, Dani LischinskiSIGGRAPH 2023 · 339 citations
Related papers
- Completing Visual Objects via Bridging Generation and SegmentationXiang Li, Yinpeng Chen, Chung-Ching Lin, Hao Chen et al.ICML 2024 · 3 citations
- SketchDream: Sketch-based Text-To-3D Generation and EditingFeng-Lin Liu, Hongbo Fu, Yu-Kun Lai, Lin GaoSIGGRAPH 2024 · 30 citations
- SmartMask: Context Aware High-Fidelity Mask Generation for Fine-grained Object Insertion and Layout ControlJaskirat Singh, Jianming Zhang, Qing Liu, Cameron Smith et al.CVPR 2024 · 7 citations
- Generative PhotomontageSean J. Liu, Nupur Kumari, Ariel Shamir, Jun-Yan ZhuCVPR 2025
- Interactive Sketch & Fill: Multiclass Sketch-to-Image TranslationArnab Ghosh, Richard Zhang, Puneet K. Dokania, Oliver Wang et al.ICCV 2019 · 148 citations
