Block and Detail: Scaffolding Sketch-to-Image Generation
Vishnu Sarukkai, Lu Yuan, Mia Tang, Maneesh Agrawala, Kayvon Fatahalian
摘要
We introduce a novel sketch-to-image tool that aligns with the iterative refinement process of artists. Our tool lets users sketch blocking strokes to coarsely represent the placement and form of objects and detail strokes to refine their shape and silhouettes. We develop a two-pass algorithm for generating high-fidelity images from such sketches at any point in the iterative process. In the first pass we use a ControlNet to generate an image that strictly follows all the strokes (blocking and detail) and in the second pass we add variation by renoising regions surrounding blocking strokes. We also present a dataset generation scheme that, when used to train a ControlNet architecture, allows regions that do not contain strokes to be interpreted as not-yet-specified regions rather than empty space. We show that this partial-sketch-aware ControlNet can generate coherent elements from partial sketches that only contain a small number of strokes. The high-fidelity images produced by our approach serve as scaffolds that can help the user adjust the shape and proportions of objects or add additional elements to the composition. We demonstrate the effectiveness of our approach with a variety of examples and evaluative comparisons. Quantitatively, evaluative user feedback indicates that novice viewers prefer the quality of images from our algorithm over a baseline Scribble ControlNet for 84% of the pairs and found our images had less distortion in 81% of the pairs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- DesignWeaver: Dimensional Scaffolding for Text-to-Image Product DesignSirui Tao, Ivan Liang, Cindy Peng, Zhiqing Wang 等CHI 2025 · 被引用 19 次
- Brickify: Enabling Expressive Design Intent Specification through Direct Manipulation on Design TokensXinyu Shi, Yinghou Wang, Ryan A. Rossi, Jian ZhaoCHI 2025 · 被引用 18 次
- GenPara: Enhancing the 3D Design Editing Process by Inferring Users' Regions of Interest with Text-Conditional Shape ParametersJiin Choi, Seung Won Lee, Kyung Hoon HyunCHI 2025 · 被引用 6 次
- DesignTrace: Exploring, Iterating and Tracking Design Alternatives with GenAIXiaohan Peng, Debanjana Haldar, Wendy E. Mackay, Janin KochCHI 2026 · 被引用 3 次
- Protosampling: Enabling Free-Form Convergence of Sampling and Prototyping through Canvas-Driven Visual AI GenerationAlicia Guo, David Ledo, George W. Fitzmaurice, Fraser AndersonCHI 2026 · 被引用 2 次
它引用的顶会 Paper10
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsChenlin Meng, Yutong He, Yang Song, Jiaming Song 等ICLR 2022 · 被引用 2,128 次
- MultiDiffusion: Fusing Diffusion Paths for Controlled Image GenerationOmer Bar-Tal, Lior Yariv, Yaron Lipman, Tali DekelICML 2023 · 被引用 575 次
- Blended Latent DiffusionOmri Avrahami, Ohad Fried, Dani LischinskiSIGGRAPH 2023 · 被引用 339 次
相关 Paper
- Completing Visual Objects via Bridging Generation and SegmentationXiang Li, Yinpeng Chen, Chung-Ching Lin, Hao Chen 等ICML 2024 · 被引用 3 次
- SketchDream: Sketch-based Text-To-3D Generation and EditingFeng-Lin Liu, Hongbo Fu, Yu-Kun Lai, Lin GaoSIGGRAPH 2024 · 被引用 30 次
- SmartMask: Context Aware High-Fidelity Mask Generation for Fine-grained Object Insertion and Layout ControlJaskirat Singh, Jianming Zhang, Qing Liu, Cameron Smith 等CVPR 2024 · 被引用 7 次
- Generative PhotomontageSean J. Liu, Nupur Kumari, Ariel Shamir, Jun-Yan ZhuCVPR 2025
- Interactive Sketch & Fill: Multiclass Sketch-to-Image TranslationArnab Ghosh, Richard Zhang, Puneet K. Dokania, Oliver Wang 等ICCV 2019 · 被引用 148 次
