SyncDiffusion: Coherent Montage via Synchronized Joint Diffusions
Yuseung Lee, Kunho Kim, Hyunjin Kim, Minhyuk Sung
Abstract
The remarkable capabilities of pretrained image diffusion models have been utilized not only for generating fixed-size images but also for creating panoramas. However, naive stitching of multiple images often results in visible seams. Recent techniques have attempted to address this issue by performing joint diffusions in multiple windows and averaging latent features in overlapping regions. However, these approaches, which focus on seamless montage generation, often yield incoherent outputs by blending different scenes within a single image. To overcome this limitation, we propose SYNCDIFFUSION, a plug-and-play module that synchronizes multiple diffusions through gradient descent from a perceptual similarity loss. Specifically, we compute the gradient of the perceptual loss using the predicted denoised images at each denoising step, providing meaningful guidance for achieving coherent montages. Our experimental results demonstrate that our method produces significantly more coherent outputs for text-guided panorama generation compared to previous methods (66.35% vs. 33.65% in our user study) while still maintaining fidelity (as assessed by GIQA) and compatibility with the input prompt (as measured by CLIP score). We further demonstrate the versatility of our method across three plug-and-play applications: layout-guided image generation, conditional image generation and 360-degree panorama generation. Our project page is at https://syncdiffusion.github.io . Figure 1: Comparison of panoramas generated with prompt "A photo of a rock concert" by Blended Latent Diffusion [1] (top), MultiDiffusion [3] (middle), and our SYNCDIFFUSION (bottom). Blended Latent Diffusion, when applied on image extrapolation, often generates visible seams and repetitive patterns. MultiDiffusion creates seamless panoramas but fails to achieve global coherence across the image. In contrast, our SYNCDIFFUSION synchronizes windows across the panorama by increasing the perceptual similarity of the denoised output predictions. This results in significantly more coherent panorama outputs. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers58
- SyncDreamer: Generating Multiview-consistent Images from a Single-view ImageYuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long et al.ICLR 2024 · 685 citations
- ScaleCrafter: Tuning-free Higher-Resolution Visual Generation with Diffusion ModelsYingqing He, Shaoshu Yang, Haoxin Chen, Xiaodong Cun et al.ICLR 2024 · 125 citations
- DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular VideosWen-Hsuan Chu, Lei Ke, Katerina FragkiadakiNeurIPS 2024 · 75 citations
- UltraPixel: Advancing Ultra High-Resolution Image Synthesis to New PeaksJingjing Ren, Wenbo Li, Haoyu Chen, Renjing Pei et al.NeurIPS 2024 · 74 citations
- Motion Guidance: Diffusion-Based Image Editing with Differentiable Motion EstimatorsDaniel Geng, Andrew OwensICLR 2024 · 46 citations
Builds on38
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- StochSync: Stochastic Diffusion Synchronization for Image Generation in Arbitrary SpacesKyeongmin Yeo, Jaihoon Kim, Minhyuk SungICLR 2025
- L-MAGIC: Language Model Assisted Generation of Images with CoherenceZhipeng Cai, Matthias Mueller, Reiner Birkl, Diana Wofk et al.CVPR 2024
- SyncNoise: Geometrically Consistent Noise Prediction for Instruction-based 3D EditingRuihuang Li, Liyi Chen, Zhengqiang Zhang, Varun Jampani et al.AAAI 2025 · 4 citations
- SPAD: Spatially Aware Multi-View DiffusersYash Kant, Aliaksandr Siarohin, Ziyi Wu, Michael Vasilkovsky et al.CVPR 2024
- MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware DiffusionShitao Tang, Fuyang Zhang, Jiacheng Chen, Peng Wang et al.NeurIPS 2023 · 249 citations
