Arbitrary-Shaped Image Generation via Spherical Neural Field Diffusion
Jiyuan Xia, Yuanshen Guan, Ruikang Xu, Zhiwei Xiong
Abstract
Existing diffusion models excel at generating diverse content, but remain confined to fixed image shapes and lack the ability to flexibly control spatial attributes such as viewpoint, field-of-view (FOV), and resolution. To fill this gap, we propose Arbitrary-Shaped Image Generation (ASIG), the first generative framework that enables precise spatial attribute control while supporting high-quality synthesis across diverse image shapes (e.g., perspective, panoramic, and fisheye). ASIG introduces two key innovations: (1) a mesh-based spherical latent diffusion to generate a complete scene representation, with seam enforcement denoising strategy to maintain semantic and spatial consistency across viewpoints; and (2) a spherical neural field to sample arbitrary regions from the scene representation with coordinate conditions, enabling distortion-free generation at flexible resolutions. To this end, ASIG enables precise control over spatial attributes within a unified framework, enabling high-quality generation across diverse image shapes. Experiments demonstrate clear improvements over prior methods specifically designed for individual shapes. Code is available at https://github.com/xjyjjy/ASIG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d2925bd-0a10-4970-a694-d86326e5716aBuilds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- SphereDiffusion: Spherical Geometry-Aware Distortion Resilient Diffusion ModelTao Wu, Xuewei Li, Zhongang Qi, Di Hu et al.AAAI 2024 · 24 citations
- 360-Degree Panorama Generation from Few Unregistered NFoV ImagesJionghao Wang, Ziyu Chen, Jun Ling, Rong Xie et al.ACM MM 2023 · 27 citations
- CubeDiff: Repurposing Diffusion-Based Image Models for Panorama GenerationNikolai Kalischek, Michael Oechsle, Fabian Manhardt, Philipp Henzler et al.ICLR 2025
- MultiDiffusion: Fusing Diffusion Paths for Controlled Image GenerationOmer Bar-Tal, Lior Yariv, Yaron Lipman, Tali DekelICML 2023 · 575 citations
- SphereDiff: Tuning-free 360° Static and Dynamic Panorama Generation via Spherical Latent RepresentationMinho Park, Taewoong Kang, Jooyeol Yun, Sungwon Hwang et al.AAAI 2026
