Latent Swap Joint Diffusion for 2D Long-Form Latent Generation
Yusheng Dai, Chenxi Wang, Chang Li, Chen Wang, Kewei Li, Jun Du, Lei Sun, Jianqing Gao, Ruoyu Wang, Jiefeng Ma
Abstract
This paper introduces Swap Forward (SaFa), a modalityagnostic and efficient method to generate seamless and coherent long spectrum and panorama using a latent swap joint diffusion process across multi-views. We first investigate spectrum aliasing problem in spectrum-based audio generation caused by existing joint diffusion methods. Through a comparative analysis of the VAE latent representation of spectra and RGB images, we identify that the failure arises from excessive suppression of high-frequency components due to the step-wise averaging operator. To address this issue, we propose Self-Loop Latent Swap, a frame-level bidirectional swap operator, applied to the overlapping region of adjacent views. Leveraging step-wise differentiated trajectories, this swap operator avoids spectrum distortion and adaptively enhances high-frequency components. Furthermore, to improve global cross-view consistency in non-overlapping regions, we introduce Reference-Guided Latent Swap, a unidirectional latent swap operator that provides a centralized reference trajectory to synchronize subview diffusions. By refining swap timing and intervals, we canachieve a balance between cross-view similarity and diversity in a feed-forward manner. Quantitative and qualitative experiments demonstrate that SaFa significantly outperforms existing joint diffusion methods and even training-based methods in audio generation using both U-Net and DiT models. It also adapts well to panorama generation, achieving comparable performance with a to speedup. The project website is available at https://swapforward.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ea0a31d0-67c4-4337-adae-ffe0719f132fCited by top-tier papers6
- Omni2Sound: Towards Unified Video-Text-to-Audio Generationyusheng dai, Zehua Chen, Yuxuan Jiang, Qiuhong Ke et al.CVPR 2026 · 12 citations
- TAG-MoE: Task-Aware Gating for Unified Generative Mixture-of-ExpertsYu Xu, Hongbin Yan, Juan Cao, Yiji Cheng et al.CVPR 2026 · 7 citations
- FreeAudio: Training-Free Timing Planning for Controllable Long-Form Text-to-Audio GenerationYuxuan Jiang, Zehua Chen, Zeqian Ju, Chang Li et al.ACM MM 2025 · 5 citations
- GeodesicNVS: Probability Density Geodesic Flow Matching for Novel View SynthesisXuqin Wang, Tao Wu, Yanfeng Zhang, Lu Liu et al.CVPR 2026 · 4 citations
- Refracting Reality: Generating Images with Realistic Transparent ObjectsYue Yin, Enze Tao, Dylan CampbellCVPR 2026 · 2 citations
Builds on28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
Related papers
- SyncDiffusion: Coherent Montage via Synchronized Joint DiffusionsYuseung Lee, Kunho Kim, Hyunjin Kim, Minhyuk SungNeurIPS 2023 · 132 citations
- Guiding a Diffusion Model by Swapping Its TokensWeijia Zhang, Yuehao Liu, Shanyan Guan, Wu Ran et al.CVPR 2026 · 2 citations
- Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent DiffusionYueming Pan, Ruoyu Feng, Qi Dai, Yuqi Wang et al.CVPR 2026 · 15 citations
- Toward Diffusible High-Dimensional Latent Spaces: A Frequency PerspectiveBolin Lai, Xudong Wang, Saketh Rambhatla, James M. Rehg et al.CVPR 2026 · 7 citations
- PanoDiffusion: 360-degree Panorama Outpainting via DiffusionTianhao Wu, Chuanxia Zheng, Tat-Jen ChamICLR 2024 · 47 citations
