SketchEvo: Leveraging Drawing Dynamics for Enhanced Image Synthesis
Zhixin Feng, Runan Yin, LAN YANG, Kaiyue Pang, Ke Li, Honggang Zhang, Yi-Zhe Song
Abstract
Sketching represents humanity's most intuitive form of visual expression -- a universal language that transcends barriers. Although recent diffusion models integrate sketches with text, they often regard the complete sketch merely as a static visual constraint, neglecting the human preference information inherently conveyed during the dynamic sketching process.This oversight leads to images that, despite technical adherence to sketches, fail to align with human aesthetic expectations. Our framework, SketchEvo, harnesses the dynamic evolution of sketches by capturing the progression from initial strokes to completed drawing. Current preference alignment techniques struggle with sketch-guided generation because the dual constraints of text and sketch create insufficiently different latent samples when using noise perturbations alone. SketchEvo addresses this through two complementary innovations: first, by leveraging sketches at different completion stages to create meaningfully divergent samples for effective aesthetic learning during training; second, through a sequence-guided rollback mechanism that applies these learned preferences during inference by balancing textual semantics with structural guidance. Extensive experiments demonstrate that these complementary approaches enable SketchEvo to deliver improved aesthetic quality while maintaining sketch fidelity, successfully generalizing to incomplete and abstract sketches throughout the drawing process.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef271c19-5bb3-486d-8bbd-b1b92b0ae92cBuilds on29
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Text to Sketch Generation with Multi-StylesTengjie Li, Shikui Tu, Lei XuNeurIPS 2025 · 1 citation
- CoProSketch: Controllable and Progressive Sketch Generation with Diffusion ModelRuohao Zhan, Yijin Li, Yisheng He, Shuo Chen et al.ACM MM 2025 · 2 citations
- One-Shot Reference-based Structure-Aware Image to Sketch SynthesisRui Yang, Honghong Yang, Li Zhao, Qin Lei et al.AAAI 2025 · 2 citations
- FlipSketch: Flipping Static Drawings to Text-Guided Sketch AnimationsHmrishav Bandyopadhyay, Yi-Zhe SongCVPR 2025
- Stroke2Sketch: Harnessing Stroke Attributes for Training-Free Sketch GenerationRui Yang, Huining Li, Yiyi Long, Xiaojun Wu et al.ICCV 2025 · 2 citations
