Lune

ICML2026Top-tier venue

Transform Trained Transformer for Accelerating Native 4K Video Generation

Jiangning Zhang, Junwei Zhu, Teng Hu, Yabiao Wang, Donghao Luo, Weijian Cao, Zhenye Gan, Xiaobin Hu, Zhucun Xue, Xiangtai Li, Chengjie Wang, Yong Liu

2026Year
3Citations

Abstract

Native 4K (2176×\times3840) video generation remains a critical challenge due to the quadratic computational explosion of full-attention as spatiotemporal resolution increases, making it difficult for models to strike a balance between efficiency and quality. This paper proposes a novel Transformer retrofit strategy termed T3 (Transform Trained Transformer) that, without altering the core architecture of full-attention pretrained models, significantly reduces compute requirements by optimizing their forward logic. Specifically, T3-Video introduces a multi-scale weight-sharing window attention mechanism and, via hierarchical blocking together with an axis-preserving full-attention design, can effect an “attention pattern” transformation of a pretrained model using only modest compute and data. Results on 4K-VBench show that T3-Video substantially outperforms existing approaches: while delivering performance improvements (+4.29↑\uparrow VQA and +0.08↑\uparrow VTC), it accelerates native 4K video generation by more than 10×\times. Demo and source code are available in #Supp.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 97cade2b-cb8d-4110-bfc9-4171cfc10f64

Builds on35

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines