Lune

CVPR2026顶会

SpeeDiff: Scalable Pixel-Anchored End-to-End Latent Diffusion Model

Bingliang Zhang, Wenda Chu, Yizhuo Li, Linjie Yang, Yisong Yue, Katherine L. Bouman, Yang Song, Qiushan Guo

出版方
2026年份

摘要

We present Scalable Pixel-anchored End-to-end Diffusion (SpeeDiff), a latent diffusion method that jointly trains the VAE and the diffusion model from scratch. In principle, joint training allows the diffusion loss gradient to directly guide the VAE encoder, encouraging the formation of a generationfriendly latent space and potentially yielding faster convergence than the conventional two-stage approach with a pretrained frozen VAE. However, a naive end-to-end implementation severely degrades performance, as unrestricted backpropagation of the diffusion loss leads to latent space collapse. Our main technical contribution is a simple yet effective Tweedie Pixel Reconstruction (TPR) loss, which provides additional pixel-level feedback by decoding a predicted clean latent from an intermediate noisy state using Tweedie's formula, thereby alleviating collapse. Furthermore, our method enables jointly scaling a fully transformerbased architecture and enhances representation alignment within the end-to-end framework. Our SpeeDiff-XL model achieves over 140ˆand 61ˆfaster training compared to Vanilla SiT and REPA, respectively, while attaining an FID of 1.50 without guidance on ImageNet 256ˆ256 generation. With a more efficient 32ˆcompressed VAE, our model further reaches an FID of 1.53 without guidance on ImageNet 512ˆ512 generation.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper33

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖