MaskSketch: Unpaired Structure-guided Masked Image Generation
Dina Bashkirova, José Lezama, Kihyuk Sohn, Kate Saenko, Irfan Essa
摘要
Recent conditional image generation methods produce images of remarkable diversity, fidelity and realism. However, the majority of these methods allow conditioning only on labels or text prompts, which limits their level of control over the generation result. In this paper, we introduce MaskSketch, an image generation method that allows spatial conditioning of the generation result using a guiding sketch as an extra conditioning signal during sampling. MaskSketch utilizes a pre-trained masked generative transformer, requiring no model training or paired supervision, and works with input sketches of different levels of abstraction. We show that intermediate self-attention maps of a masked generative transformer encode important structural information of the input image, such as scene layout and object shape, and we propose a novel sampling method based on this observation to enable structure-guided generation. Our results show that MaskSketch achieves high image realism and fidelity to the guiding structure. Evaluated on standard benchmark datasets, MaskSketch outperforms stateof-the-art methods for sketch-to-image translation, as well as unpaired image-to-image translation approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- BlockFusion: Expandable 3D Scene Generation using Latent Tri-plane ExtrapolationZhennan Wu, Yang Li, Han Yan, Taizhang Shang 等SIGGRAPH 2024 · 被引用 28 次
- Slight Corruption in Pre-training Data Makes Better Diffusion ModelsHao Chen, Yujin Han, Diganta Misra, Xiang Li 等NeurIPS 2024 · 被引用 14 次
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image EditingWei Chow, Linfeng Li, Lingdong Kong, Zefeng Li 等CVPR 2026 · 被引用 14 次
- MaskINT: Video Editing via Interpolative Non-autoregressive Masked TransformersHaoyu Ma, Shahin Mahdizadehaghdam, Bichen Wu, Zhipeng Fan 等CVPR 2024 · 被引用 3 次
- RobuSTereo: Robust Zero-Shot Stereo Matching under Adverse WeatherYuran Wang, Yingping Liang, Yutao Hu, Ying FuICCV 2025 · 被引用 3 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- Sketch-Guided Text-to-Image Diffusion ModelsAndrey Voynov, Kfir Aberman, Daniel Cohen-OrSIGGRAPH 2023 · 被引用 168 次
- Interactive Sketch & Fill: Multiclass Sketch-to-Image TranslationArnab Ghosh, Richard Zhang, Puneet K. Dokania, Oliver Wang 等ICCV 2019 · 被引用 148 次
- Guided Image-to-Image Translation With Bi-Directional Feature TransformationBadour Albahar, Jia-Bin HuangICCV 2019 · 被引用 102 次
- Text to Sketch Generation with Multi-StylesTengjie Li, Shikui Tu, Lei XuNeurIPS 2025 · 被引用 1 次
- Modulating Pretrained Diffusion Models for Multimodal Image SynthesisCusuh Ham, James Hays, Jingwan Lu, Krishna Kumar Singh 等SIGGRAPH 2023 · 被引用 17 次
