Lune

CVPR2026Top-tier venue

Dual-Granularity Memory for Efficient Video Generation

Hongjun Wang, Lin Liu, Jianguo Li, Tao Lin

2026Year

Abstract

Recurrent architectures offer significant efficiency advantages over attention-based transformers for video generation, particularly for long sequences. However, chunked processing in these models creates temporal discontinuities that degrade long-range consistency. We introduce two complementary memory mechanisms that address this problem at different granularities. Context Memory maintains persistent global context within processing chunks through learnable sink columns and boundary buffers, adding only 150K parameters (less than 0.1% overhead). Latent Context-as-Memory (LCaM) extends memory across video segments by storing and retrieving historical latent embeddings, enabling cross-segment consistency without camera annotations or frame reconstruction. Applied to Generalized Spatial-Temporal Propagation Networks (GSTPN) distilled from WanVideo-1.3B, our dualmemory approach achieves 1.54× faster inference than attention-based transformers while maintaining competitive visual quality. LCaM is particularly suited to knowledge distillation settings where only pre-extracted latent embeddings are available. Extensive experiments and ablations validate the effectiveness of memory-augmented recurrent architectures for practical long-context video generation.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Builds on35

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines