SketchingReality: From Freehand Scene Sketches to Photorealistic Images
Ahmed Bourouis, Mikhail Bessmeltsev, Yulia Gryaditskaya
Abstract
Recent years have witnessed remarkable progress in generative AI, with natural language emerging as the most common conditioning input. As underlying models grow more powerful, researchers are exploring increasingly diverse conditioning signals -- such as depth maps, edge maps, camera parameters, and reference images -- to give users finer control over generation. Among different modalities, sketches constitute a natural and long-standing form of human communication, enabling rapid expression of visual concepts. Yet algorithms that effectively handle true freehand sketches -- with their inherent abstraction and distortions -- remain largely unexplored. In this work, we distinguish between edge maps, often regarded as “sketches” in the literature, and genuine freehand sketches. We pursue the challenging goal of balancing photorealism with sketch adherence when generating images from freehand input. A key obstacle is the absence of ground-truth, pixel-aligned images: by their nature, freehand sketches do not have a single correct alignment. To address this, we propose a modulation-based approach that prioritizes semantic interpretation of the sketch over strict adherence to individual edge positions. We further introduce a novel loss that enables training on freehand sketches without requiring ground-truth pixel-aligned images. We show that our method outperforms existing approaches in both semantic alignment with freehand sketch inputs and in the realism and overall quality of the generated images.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b5532f25-1145-454e-b1eb-a9fb66ffc8deBuilds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion ModelsChong Mou, Xintao Wang, Liangbin Xie, Yanze Wu et al.AAAI 2024 · 1,641 citations
- Composer: Creative and Controllable Image Synthesis with Composable ConditionsLianghua Huang, Di Chen, Yu Liu, Yujun Shen et al.ICML 2023 · 371 citations
Related papers
- Picture that Sketch: Photorealistic Image Generation from Abstract SketchesSubhadeep Koley, Ayan Kumar Bhunia, Aneeshan Sain, Pinaki Nath Chowdhury et al.CVPR 2023
- SketchyCOCO: Image Generation From Freehand Scene SketchesChengying Gao, Qi Liu, Qi Xu, Limin Wang et al.CVPR 2020
- DeepFaceDrawing: deep generation of face images from sketchesShu-Yu Chen, Wanchao Su, Lin Gao, Shihong Xia et al.SIGGRAPH 2020 · 145 citations
- Stroke2Sketch: Harnessing Stroke Attributes for Training-Free Sketch GenerationRui Yang, Huining Li, Yiyi Long, Xiaojun Wu et al.ICCV 2025 · 2 citations
- It's All About Your Sketch: Democratising Sketch Control in Diffusion ModelsSubhadeep Koley, Ayan Kumar Bhunia, Deeptanshu Sekhri, Aneeshan Sain et al.CVPR 2024
