Guided and Variance-Corrected Fusion with One-shot Style Alignment for Large-Content Image Generation
Shoukun Sun, Min Xian, Tiankai Yao, Fei Xu, Luca Capriotti
摘要
Producing large images using small diffusion models is gaining increasing popularity, as the cost of training large models could be prohibitive. A common approach involves jointly generating a series of overlapped image patches and obtaining large images by merging adjacent patches. However, results from existing methods often exhibit noticeable artifacts, e.g., seams and inconsistent objects and styles. To address the issues, we proposed Guided Fusion (GF), which mitigates the negative impact from distant image regions by applying a weighted average to the overlapping regions. Moreover, we proposed Variance-Corrected Fusion (VCF), which corrects data variance at post-averaging, generating more accurate fusion for the Denoising Diffusion Probabilistic Model. Furthermore, we proposed a one-shot Style Alignment (SA), which generates a coherent style for large images by adjusting the initial input noise without adding extra computational burden. Extensive experiments demonstrated that the proposed fusion methods improved the quality of the generated image significantly. The proposed method can be widely applied as a plug-and-play module to enhance other fusion-based methods for large image generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
相关 Paper
- Conditional Controllable Image FusionBing Cao, Xingxin Xu, Pengfei Zhu, Qilong Wang 等NeurIPS 2024 · 被引用 29 次
- VMDiff: Visual Mixing Diffusion for Limitless Cross-Object SynthesisZeren Xiong, Yue Yu, Ze-dong Zhang, Shuo Chen 等ICLR 2026 · 被引用 1 次
- InfinityGAN: Towards Infinite-Pixel Image SynthesisChieh Hubert Lin, Hsin-Ying Lee, Yen-Chi Cheng, Sergey Tulyakov 等ICLR 2022 · 被引用 84 次
- InstantAS: Minimum Coverage Sampling for Arbitrary-Size Image GenerationChangshuo Wang, Mingzhe Yu, Lei Wu, Lei Meng 等ACM MM 2024
- Efficient and Training-Free Single-Image Diffusion ModelsHaojun Qiu, Kiriakos N. Kutulakos, David B. LindellCVPR 2026 · 被引用 1 次
