Lune

ICLR2024顶会

InstaFlow: One Step is Enough for High-Quality Diffusion-Based Text-to-Image Generation

Xingchao Liu, Xiwen Zhang, Jianzhu Ma, Jian Peng, Qiang Liu

2024年份
358被引次数
218顶会引用

摘要

Diffusion models have revolutionized text-to-image generation with its exceptional quality and creativity. However, its multi-step sampling process is known to be slow, often requiring tens of inference steps to obtain satisfactory results. Previous attempts to improve its sampling speed and reduce computational costs through distillation have been unsuccessful in achieving a functional one-step model. In this paper, we explore a recent method called Rectified Flow [45; 43], which, thus far, has only been applied to small datasets. The core of Rectified Flow lies in its reflow procedure, which straightens the trajectories of probability flows, refines the coupling between noises and images, and facilitates the distillation process with student models. We propose a novel text-conditioned pipeline to turn Stable Diffusion (SD) into an ultra-fast one-step model, in which we find reflow plays a critical role in improving the assignment between noises and images. Leveraging our new pipeline, we create, to the best of our knowledge, the first one-step diffusion-based text-to-image generator with SD-level image quality, achieving an FID (Fréchet Inception Distance) of 23.3 on MS COCO 2017-5k, surpassing the previous state-of-the-art technique, progressive distillation [58] , by a significant margin (37.2 → 23.3 in FID). By utilizing an expanded network with 1.7B parameters, we further improve the FID to 22.4. We call our onestep models InstaFlow. On MS COCO 2014-30k, InstaFlow yields an FID of 13.1 in just 0.09 second, the best in ≤ 0.1 second regime, outperforming the recent StyleGAN-T [73] (13.9 in 0.1 second). Notably, the training of InstaFlow only costs 199 A100 GPU days. Codes and pre-trained models are available at github.com/gnobitab/InstaFlow. Figure 1: InstaFlow is a high-quality one-step text-to-image model derived from Stable Diffusion [70]. Within 0.1 second, it generates images with similar FID as StyleGAN-T [73] on MS COCO 2014. The whole finetuning process to yield InstaFlow is pure supervised learning and costs only 199 A100 GPU days.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper218

问问它们各自怎么用它

它引用的顶会 Paper67

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖