World-Shaper: A Unified Framework for 360° Panoramic Editing
Dong Liang, yuhao liu, Jinyuan Jia, Youjun Zhao, Rynson W Lau
Abstract
Being able to edit panoramic images is crucial for creating realistic 360° visual experiences. However, existing perspective-based image editing methods fail to model the spatial structure of panoramas. Conventional cube-map decompositions attempt to overcome this problem but inevitably break global consistency due to their mismatch with spherical geometry. Motivated by this insight, we reformulate panoramic editing directly in the equirectangular projection (ERP) domain and present World Shaper, a unified geometry-aware framework that supports five distinct editing operations within a single ERP-native representation. To address the latitude-dependent geometric distortion inherent in ERP, we introduce a geometry-aware learning strategy comprising distortion-aware attention modulation (DAAM), which steers cross-attention with latitude-dependent strength at the feature level; layered shape loss (LSL), which enforces per-object geometric supervision at the output level; and progressive curriculum training to internalize panoramic priors. To overcome the scarcity of paired panoramic editing data, we train a dedicated ERP-native controllable generator that synthesizes objects directly in the equirectangular domain under user-defined conditions, enabling scalable paired data construction for learning diverse editing behaviors. Extensive experiments on our new benchmark, PEBench, demonstrate that World Shaper achieves superior geometric consistency, editing fidelity, and text controllability compared to state-of-the-art methods, enabling coherent and flexible 360° visual world creation with unified editing control. Code, models, and data are available at the project page: https://world-shaper-project.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f9dd6ea9-8200-43ad-b174-33d5491382bcBuilds on37
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- PanoSwin: a Pano-style Swin Transformer for Panorama UnderstandingZhixin Ling, Zhen Xing, Xiangdong Zhou, Manliang Cao et al.CVPR 2023
- SE360: Semantic Edit in 360° Panoramas via Hierarchical Data ConstructionHaoyi Zhong, Fang-Lue Zhang, Andrew Chalmers, Taehyun RheeAAAI 2026
- Conditional Panoramic Image Generation via Masked Autoregressive ModelingChaoyang Wang, Xiangtai Li, Lu Qi, Xiaofan Lin et al.NeurIPS 2025 · 11 citations
- PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video DiffusionYuyang Yin, Hao-Xiang Guo, Fangfu Liu, Mengyu Wang et al.ICML 2026 · 3 citations
- SphereDiffusion: Spherical Geometry-Aware Distortion Resilient Diffusion ModelTao Wu, Xuewei Li, Zhongang Qi, Di Hu et al.AAAI 2024 · 24 citations
