Diffusion-4K: Ultra-High-Resolution Image Synthesis with Latent Diffusion Models
Jinjin Zhang, Qiuyu Huang, Junjie Liu, Xiefan Guo, Di Huang
摘要
In this paper, we present Diffusion-4K, a novel framework for direct ultra-high-resolution image synthesis using text-to-image diffusion models. The core advancements include: (1) Aesthetic-4K Benchmark: addressing the absence of a publicly available 4K image synthesis dataset, we construct Aesthetic-4K, a comprehensive benchmark for ultra-high-resolution image generation. We curated a high-quality 4K dataset with carefully selected images and captions generated by GPT-4o. Additionally, we introduce GLCM Score and Compression Ratio metrics to evaluate fine details, combined with holistic measures such as FID, Aesthetics and CLIPScore for a comprehensive assessment of ultra-high-resolution images. (2) Wavelet-based Fine-tuning: we propose a wavelet-based fine-tuning approach for direct training with photorealistic 4K images, applicable to various latent diffusion models, demonstrating its effectiveness in synthesizing highly detailed 4K images. Consequently, Diffusion-4K achieves impressive performance in high-quality image synthesis and text prompt adherence, especially when powered by modern large-scale diffusion models (e.g., SD3-2B and Flux-12B). Extensive experimental results from our benchmark demonstrate the superiority of Diffusion-4K in ultra-high-resolution image synthesis. Code is available at https://github.com/zhang0jhon/diffusion-4k.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- 4KAgent: Agentic Any Image to 4K Super-ResolutionYushen Zuo, Qi Zheng, Mingyang Wu, Xinrui Jiang 等NeurIPS 2025 · 被引用 51 次
- Fast-Slow Thinking GRPO for Large Vision-Language Model ReasoningWenyi Xiao, Leilei GanNeurIPS 2025 · 被引用 34 次
- UltraHR-100K: Enhancing UHR Image Synthesis with A Large-Scale High-Quality DatasetChen Zhao, En Ci, Yunzhe Xu, Tiehan Fan 等NeurIPS 2025 · 被引用 24 次
- HiFlow: Training-free High-Resolution Image Generation with Flow-Aligned GuidanceJiazi Bu, Pengyang Ling, Yujie Zhou, Pan Zhang 等NeurIPS 2025 · 被引用 21 次
- DyPE: Dynamic Position Extrapolation for Ultra High Resolution DiffusionNoam Issachar, Guy Yariv, Sagie Benaim, Yossi Adi 等ICML 2026 · 被引用 14 次
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
相关 Paper
- UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect RatiosTian Ye, Song Fei, Lei ZhuCVPR 2026 · 被引用 9 次
- Latent Wavelet Diffusion For Ultra High-Resolution Image SynthesisLuigi Sigillo, Shengfeng He, Danilo ComminielloICLR 2026 · 被引用 8 次
- RealUHR: Harnessing Patch-Cascade Flows for Photorealistic Ultra-High-Resolution SynthesisYongsheng Yu, Haitian Zheng, Zhe Lin, Connelly Barnes 等AAAI 2026
- Ultra-Resolution Adaptation with EaseRuonan Yu, Songhua Liu, Zhenxiong Tan, Xinchao WangICML 2025
- DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?Qirui Jiao, Daoyuan Chen, Yilun Huang, Xika Lin 等ICML 2026
