Generative Powers of Ten
Xiaojuan Wang, Janne Kontkanen, Brian Curless, Steven M. Seitz, Ira Kemelmacher-Shlizerman, Ben Mildenhall, Pratul P. Srinivasan, Dor Verbin, Aleksander Holynski
摘要
We present a method that uses a text-to-image model to generate consistent content across multiple image scales, enabling extreme semantic zooms into a scene, e.g. ranging from a wide-angle landscape view of a forest to a macro shot of an insect sitting on one of the tree branches. We achieve this through a joint multi-scale diffusion sampling approach that encourages consistency across different scales while preserving the integrity of each individual sampling process. Since each generated scale is guided by a different text prompt, our method enables deeper levels of zoom than traditional super-resolution methods that may struggle to create new contextual structure at vastly different scales. We compare our method qualitatively with alter-native techniques in image super-resolution and outpainting, and show that our method is most effective at generating consistent multi-scale content.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- SyncTweedies: A General Generative Framework Based on Synchronized DiffusionsJaihoon Kim, Juil Koo, Kyeongmin Yeo, Minhyuk SungNeurIPS 2024 · 被引用 27 次
- Chain-of-Zoom: Extreme Super-Resolution via Scale Autoregression and Preference AlignmentBryan Sangwoo Kim, Jeongsol Kim, Jong Chul YeNeurIPS 2025 · 被引用 10 次
- D-VST: Diffusion Transformer for Pathology-Correct Tone-Controllable Cross-Dye Virtual Staining of Whole Slide ImagesShurong Yang, Dong Wei, Yihuang Hu, Qiong Peng 等NeurIPS 2025 · 被引用 2 次
- WonderZoom: Multi-Scale 3D World GenerationJin Cao, Hong-Xing Yu, Jiajun WuCVPR 2026 · 被引用 1 次
- LookingGlass: Generative Anamorphoses via Laplacian Pyramid WarpingPascal Chang, Sergio Sancho, Jingwei Tang, Markus Gross 等CVPR 2025
它引用的顶会 Paper17
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- Text-Guided Explorable Image Super-ResolutionKanchana Vaishnavi Gandikota, Paramanand ChandramouliCVPR 2024
- Ultra High-Resolution Image Inpainting with Patch-Based Content Consistency AdapterJianhui Zhang, Sheng Cheng, Qirui Sun, Jia Liu 等ICCV 2025 · 被引用 1 次
- Nested Scale-Editing for Conditional Image SynthesisLingzhi Zhang, Jiancong Wang, Yinshuang Xu, Jie Min 等CVPR 2020
- MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware DiffusionShitao Tang, Fuyang Zhang, Jiacheng Chen, Peng Wang 等NeurIPS 2023 · 被引用 249 次
- Text2Tex: Text-driven Texture Synthesis via Diffusion ModelsDave Zhenyu Chen, Yawar Siddiqui, Hsin-Ying Lee, Sergey Tulyakov 等ICCV 2023 · 被引用 262 次
