Lune

CVPR2026Top-tier venue

SpeeDiff: Scalable Pixel-Anchored End-to-End Latent Diffusion Model

Bingliang Zhang, Wenda Chu, Yizhuo Li, Linjie Yang, Yisong Yue, Katherine L. Bouman, Yang Song, Qiushan Guo

2026Year

Abstract

We present Scalable Pixel-anchored End-to-end Diffusion (SpeeDiff), a latent diffusion method that jointly trains the VAE and the diffusion model from scratch. In principle, joint training allows the diffusion loss gradient to directly guide the VAE encoder, encouraging the formation of a generationfriendly latent space and potentially yielding faster convergence than the conventional two-stage approach with a pretrained frozen VAE. However, a naive end-to-end implementation severely degrades performance, as unrestricted backpropagation of the diffusion loss leads to latent space collapse. Our main technical contribution is a simple yet effective Tweedie Pixel Reconstruction (TPR) loss, which provides additional pixel-level feedback by decoding a predicted clean latent from an intermediate noisy state using Tweedie's formula, thereby alleviating collapse. Furthermore, our method enables jointly scaling a fully transformerbased architecture and enhances representation alignment within the end-to-end framework. Our SpeeDiff-XL model achieves over 140ˆand 61ˆfaster training compared to Vanilla SiT and REPA, respectively, while attaining an FID of 1.50 without guidance on ImageNet 256ˆ256 generation. With a more efficient 32ˆcompressed VAE, our model further reaches an FID of 1.53 without guidance on ImageNet 512ˆ512 generation.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext a28ec2e0-6352-42e0-8a6f-df87c2375f16

Builds on33

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines