DiT360: High-Fidelity Panoramic Image Generation via Hybrid Training
Haoran Feng, Dizhe Zhang, Xiangtai Li, Bo Du, Lu Qi
摘要
In this work, we propose DiT360, a DiT-based framework that performs hybrid training on perspective and panoramic data for panoramic image generation. We attribute the main challenges in preserving geometric fidelity and photorealism to the scarcity of large-scale, highquality real-world panoramic data, in contrast to prior methods that emphasize model design. Basically, DiT360 has several key modules for inter-domain transformation and intra-domain augmentation, applied at both the pre-
VAE image level and the post-VAE token level. At the image level, we incorporate cross-domain knowledge through perspective image guidance and panoramic refinement, which enhance perceptual quality while regularizing diversity and photorealism. At the token level, hybrid supervision is applied across multiple modules, which include circular padding for boundary continuity, yaw loss for rotational robustness, and cube loss for distortion awareness. Extensive experiments on text-to-panorama, inpainting, and outpainting tasks demonstrate that our method achieves better boundary consistency and image fidelity across eleven quantitative metrics. Our code is available This CVPR paper is the Open Access version, provided by the Computer Vision Foundation.
Except for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Depth Any Panoramas: A Foundation Model for Panoramic Depth EstimationXin Lin, Meixi Song, Dizhe Zhang, Wenxuan Lu 等CVPR 2026 · 被引用 27 次
- CineScene: Implicit 3D as Effective Scene Representation for Cinematic Video GenerationKaiyi Huang, Yukun Huang, Yu Li, Jianhong Bai 等CVPR 2026 · 被引用 7 次
- Pano360: Perspective to Panoramic Vision with Geometric ConsistencyZhengdong Zhu, Weiyi Xue, Zuyuan Yang, Wenlve Zhou 等CVPR 2026
- MTPano: Multi-Task Panoramic Scene Understanding via Label-Free Integration of Dense Prediction PriorsJingdong Zhang, Xiaohang Zhan, Lingzhi Zhang, Yizhou Wang 等SIGGRAPH 2026
它引用的顶会 Paper43
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- PanoDiT: Panoramic Videos Generation with Diffusion TransformerMuyang Zhang, Yuzhi Chen, Rongtao Xu, Changwei Wang 等AAAI 2025 · 被引用 6 次
- Dream360: Diverse and Immersive Outdoor Virtual Scene Creation via Transformer-Based 360° Image OutpaintingHao Ai, Zidong Cao, Haonan Lu, Chen Chen 等IEEE VR 2024 · 被引用 12 次
- 360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion ModelQian Wang, Weiqi Li, Chong Mou, Xinhua Cheng 等CVPR 2024 · 被引用 23 次
- SE360: Semantic Edit in 360° Panoramas via Hierarchical Data ConstructionHaoyi Zhong, Fang-Lue Zhang, Andrew Chalmers, Taehyun RheeAAAI 2026
- Pano3DComposer: Feed-Forward Compositional 3D Scene Generation from Single Panoramic ImageZidian Qiu, Ancong WuCVPR 2026 · 被引用 1 次
