SketchEvo: Leveraging Drawing Dynamics for Enhanced Image Synthesis
Zhixin Feng, Runan Yin, LAN YANG, Kaiyue Pang, Ke Li, Honggang Zhang, Yi-Zhe Song
摘要
Sketching represents humanity's most intuitive form of visual expression -- a universal language that transcends barriers. Although recent diffusion models integrate sketches with text, they often regard the complete sketch merely as a static visual constraint, neglecting the human preference information inherently conveyed during the dynamic sketching process.This oversight leads to images that, despite technical adherence to sketches, fail to align with human aesthetic expectations. Our framework, SketchEvo, harnesses the dynamic evolution of sketches by capturing the progression from initial strokes to completed drawing. Current preference alignment techniques struggle with sketch-guided generation because the dual constraints of text and sketch create insufficiently different latent samples when using noise perturbations alone. SketchEvo addresses this through two complementary innovations: first, by leveraging sketches at different completion stages to create meaningfully divergent samples for effective aesthetic learning during training; second, through a sequence-guided rollback mechanism that applies these learned preferences during inference by balancing textual semantics with structural guidance. Extensive experiments demonstrate that these complementary approaches enable SketchEvo to deliver improved aesthetic quality while maintaining sketch fidelity, successfully generalizing to incomplete and abstract sketches throughout the drawing process.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper29
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- Text to Sketch Generation with Multi-StylesTengjie Li, Shikui Tu, Lei XuNeurIPS 2025 · 被引用 1 次
- CoProSketch: Controllable and Progressive Sketch Generation with Diffusion ModelRuohao Zhan, Yijin Li, Yisheng He, Shuo Chen 等ACM MM 2025 · 被引用 2 次
- One-Shot Reference-based Structure-Aware Image to Sketch SynthesisRui Yang, Honghong Yang, Li Zhao, Qin Lei 等AAAI 2025 · 被引用 2 次
- FlipSketch: Flipping Static Drawings to Text-Guided Sketch AnimationsHmrishav Bandyopadhyay, Yi-Zhe SongCVPR 2025
- Stroke2Sketch: Harnessing Stroke Attributes for Training-Free Sketch GenerationRui Yang, Huining Li, Yiyi Long, Xiaojun Wu 等ICCV 2025 · 被引用 2 次
