SphereDiffusion: Spherical Geometry-Aware Distortion Resilient Diffusion Model
Tao Wu, Xuewei Li, Zhongang Qi, Di Hu, Xintao Wang, Ying Shan, Xi Li
Abstract
Controllable spherical panoramic image generation holds substantial applicative potential across a variety of domains. However, it remains a challenging task due to the inherent spherical distortion and geometry characteristics, resulting in low-quality content generation. In this paper, we introduce a novel framework of SphereDiffusion to address these unique challenges, for better generating high-quality and precisely controllable spherical panoramic images. For the spherical distortion characteristic, we embed the semantics of the distorted object with text encoding, then explicitly construct the relationship with text-object correspondence to better use the pre-trained knowledge of the planar images. Meanwhile, we employ a deformable technique to mitigate the semantic deviation in latent space caused by spherical distortion. For the spherical geometry characteristic, in virtue of spherical rotation invariance, we improve the data diversity and optimization objectives in the training process, enabling the model to better learn the spherical geometry characteristic. Furthermore, we enhance the denoising process of the diffusion model, enabling it to effectively use the learned geometric characteristic to ensure the boundary continuity of the generated images. With these specific techniques, experiments on Structured3D dataset show that SphereDiffusion significantly improves the quality of controllable spherical image generation and relatively reduces around 35% FID on average.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ad2703e-4a1a-4fb2-a64f-a0f5a8b0d68bCited by top-tier papers14
- DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image GenerationZehong Ma, Longhui Wei, Shuai Wang, Shiliang Zhang et al.CVPR 2026 · 59 citations
- MagCache: Fast Video Generation with Magnitude-Aware CacheZehong Ma, Longhui Wei, Feng Wang, Shiliang Zhang et al.NeurIPS 2025 · 41 citations
- Conditional Panoramic Image Generation via Masked Autoregressive ModelingChaoyang Wang, Xiangtai Li, Lu Qi, Xiaofan Lin et al.NeurIPS 2025 · 11 citations
- Target-Driven Distillation: Consistency Distillation with Target Timestep Selection and Decoupled GuidanceCunzheng Wang, Ziyuan Guo, Yuxuan Duan, Huaxia Li et al.AAAI 2025 · 9 citations
- RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera ControlTeng Li, Guangcong Zheng, Rui Jiang, Shuigen Zhan et al.ICCV 2025 · 5 citations
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Spherical Manifold Guided Diffusion Model for Panoramic Image GenerationXiancheng Sun, Mai Xu, Shengxi Li, Senmao Ma et al.CVPR 2025
- DiffPano: Scalable and Consistent Text to Panorama Generation with Spherical Epipolar-Aware DiffusionWeicai Ye, Chenhao Ji, Zheng Chen, Junyao Gao et al.NeurIPS 2024 · 45 citations
- SphereDiff: Tuning-free 360° Static and Dynamic Panorama Generation via Spherical Latent RepresentationMinho Park, Taewoong Kang, Jooyeol Yun, Sungwon Hwang et al.AAAI 2026
- Spherical-Nested Diffusion Model for Panoramic Image OutpaintingXiancheng Sun, Senmao Ma, Shengxi Li, Mai Xu et al.ICML 2025
- Arbitrary-Shaped Image Generation via Spherical Neural Field DiffusionJiyuan Xia, Yuanshen Guan, Ruikang Xu, Zhiwei XiongICLR 2026
