FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers
Minguk Kang, Suha Kwak
摘要
Recent progress in video generation has shifted large-scale models from convolutional architectures to Diffusion Transformers (DiT), yet latent-to-pixel video decoders remain predominantly convolutional. These decoders rely on heavy 3D convolutions, which slow down streaming generation and require spatial–temporal tiling to handle high-resolution or long-duration outputs. We introduce FlashDecoder, the first Transformer-based latent-to-pixel video decoder designed for streaming. FlashDecoder processes video latents frame-by-frame during both training and inference, applying bidirectional spatial attention within each frame while maintaining causal temporal dependencies through a rolling KV cache. Crucially, causality is enforced by sequential frame processing rather than explicit attention masks, enabling the use of memory-efficient bidirectional attention kernels throughout. This unified streaming approach ensures constant per-frame computation and bounded memory via a fixed-size KV cache with automatic eviction of older frames, enabling stable training at resolutions up to 720p. Integrated into the Wan2.2 video VAE, FlashDecoder matches the reconstruction quality of the convolutional decoder (PSNR 38.38 vs. 38.29; LPIPS 0.046 vs. 0.039) while decoding up to 4x faster—139 FPS at 480p and 69.6 FPS at 720p—achieving real-time high-resolution video decoding on a single H100 GPU.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper36
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
相关 Paper
- Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video GenerationLunjie Zhu, Yushi Huang, Xingtong Ge, Yufei Xue 等ICML 2026
- Faster Video Diffusion with Trainable Sparse AttentionPeiyuan Zhang, Yongqi Chen, Haofeng Huang, Will Lin 等NeurIPS 2025 · 被引用 6 次
- Veda: Scalable Video Diffusion via Distilled Sparse AttentionShihao Han, Hao Yang, Xiaofeng Mei, Xinting Hu 等ICML 2026 · 被引用 1 次
- FlashVSR: Towards Real-time Diffusion-Based Streaming Video Super ResolutionJunhao Zhuang, Shi Guo, Xin Cai, Xiaohui Li 等CVPR 2026 · 被引用 42 次
- Streaming Visual Geometry TransformerDong Zhuo, Wenzhao Zheng, Jiahe Guo, Yuqi Wu 等ICLR 2026 · 被引用 109 次
