Spherical Manifold Guided Diffusion Model for Panoramic Image Generation
Xiancheng Sun, Mai Xu, Shengxi Li, Senmao Ma, Xin Deng, Lai Jiang, Gang Shen
Abstract
Panoramic image essentially acts as a pivotal role in emerging virtual reality and augmented reality scenarios; however, the generation of panoramic images are essentially challenging due to the intrinsic spherical geometry and spherical distortions caused by equirectangular projection (ERP). To address this, we start from the very basics of S2manifold inherent to panoramic images, and propose a novel spherical manifold convolution (SMConv) on S2manifold. Based on the SMConv operation, we propose a spherical manifold guided diffusion (SMGD) model for text-conditioned panoramic image generation, which can well accommodate the spherical geometry during generation. We further develop a novel evaluation method by calculating grouped Fréchet inception distance (FID) on cube-map projections, which can well reflect the quality of generated panoramic images, compared to existing methods that randomly crop ERP-distorted content. Experiment results demonstrate that our SMGD model achieves the state-of-the-art generation quality and accuracy, whilst retaining the shortest sampling time in the text-conditioned panoramic image generation task. Codes are publicly available at https://github.com/chronos123/SMGD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 955d76b9-5239-425c-ae67-7935666b6919Cited by top-tier papers3
- DiT360: High-Fidelity Panoramic Image Generation via Hybrid TrainingHaoran Feng, Dizhe Zhang, Xiangtai Li, Bo Du et al.CVPR 2026 · 27 citations
- Arbitrary-Shaped Image Generation via Spherical Neural Field DiffusionJiyuan Xia, Yuanshen Guan, Ruikang Xu, Zhiwei XiongICLR 2026
- World-Shaper: A Unified Framework for 360° Panoramic EditingDong Liang, yuhao liu, Jinyuan Jia, Youjun Zhao et al.ICML 2026
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- SphereDiffusion: Spherical Geometry-Aware Distortion Resilient Diffusion ModelTao Wu, Xuewei Li, Zhongang Qi, Di Hu et al.AAAI 2024 · 24 citations
- SphereDiff: Tuning-free 360° Static and Dynamic Panorama Generation via Spherical Latent RepresentationMinho Park, Taewoong Kang, Jooyeol Yun, Sungwon Hwang et al.AAAI 2026
- Conditional Panoramic Image Generation via Masked Autoregressive ModelingChaoyang Wang, Xiangtai Li, Lu Qi, Xiaofan Lin et al.NeurIPS 2025 · 11 citations
- CubeDiff: Repurposing Diffusion-Based Image Models for Panorama GenerationNikolai Kalischek, Michael Oechsle, Fabian Manhardt, Philipp Henzler et al.ICLR 2025
- Spherical Pseudo-Cylindrical Representation for Omnidirectional Image Super-resolutionQing Cai, Mu Li, Dongwei Ren, Jun Lyu et al.AAAI 2024 · 11 citations
