Generative Powers of Ten
Xiaojuan Wang, Janne Kontkanen, Brian Curless, Steven M. Seitz, Ira Kemelmacher-Shlizerman, Ben Mildenhall, Pratul P. Srinivasan, Dor Verbin, Aleksander Holynski
Abstract
We present a method that uses a text-to-image model to generate consistent content across multiple image scales, enabling extreme semantic zooms into a scene, e.g. ranging from a wide-angle landscape view of a forest to a macro shot of an insect sitting on one of the tree branches. We achieve this through a joint multi-scale diffusion sampling approach that encourages consistency across different scales while preserving the integrity of each individual sampling process. Since each generated scale is guided by a different text prompt, our method enables deeper levels of zoom than traditional super-resolution methods that may struggle to create new contextual structure at vastly different scales. We compare our method qualitatively with alter-native techniques in image super-resolution and outpainting, and show that our method is most effective at generating consistent multi-scale content.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85d7b51b-12ad-44dc-ba76-ab3fbb93437dCited by top-tier papers7
- SyncTweedies: A General Generative Framework Based on Synchronized DiffusionsJaihoon Kim, Juil Koo, Kyeongmin Yeo, Minhyuk SungNeurIPS 2024 · 27 citations
- Chain-of-Zoom: Extreme Super-Resolution via Scale Autoregression and Preference AlignmentBryan Sangwoo Kim, Jeongsol Kim, Jong Chul YeNeurIPS 2025 · 10 citations
- D-VST: Diffusion Transformer for Pathology-Correct Tone-Controllable Cross-Dye Virtual Staining of Whole Slide ImagesShurong Yang, Dong Wei, Yihuang Hu, Qiong Peng et al.NeurIPS 2025 · 2 citations
- WonderZoom: Multi-Scale 3D World GenerationJin Cao, Hong-Xing Yu, Jiajun WuCVPR 2026 · 1 citation
- LookingGlass: Generative Anamorphoses via Laplacian Pyramid WarpingPascal Chang, Sergio Sancho, Jingwei Tang, Markus Gross et al.CVPR 2025
Builds on17
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Text-Guided Explorable Image Super-ResolutionKanchana Vaishnavi Gandikota, Paramanand ChandramouliCVPR 2024
- Ultra High-Resolution Image Inpainting with Patch-Based Content Consistency AdapterJianhui Zhang, Sheng Cheng, Qirui Sun, Jia Liu et al.ICCV 2025 · 1 citation
- Nested Scale-Editing for Conditional Image SynthesisLingzhi Zhang, Jiancong Wang, Yinshuang Xu, Jie Min et al.CVPR 2020
- MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware DiffusionShitao Tang, Fuyang Zhang, Jiacheng Chen, Peng Wang et al.NeurIPS 2023 · 249 citations
- Text2Tex: Text-driven Texture Synthesis via Diffusion ModelsDave Zhenyu Chen, Yawar Siddiqui, Hsin-Ying Lee, Sergey Tulyakov et al.ICCV 2023 · 262 citations
