Lune

AAAI2026Top-tier venue

FreeMem: Enhancing Consistency in Long Video Generation via Tuning-Free Memory

Jibin Peng, Di Lin, Zhecheng Xu, Haoran Lu, Ruonan Liu, Wuyuan Xie, Miaohui Wang, Lingyu Liang, Yi Wang, Qing Guo

2026Year

Abstract

Text-to-Video (T2V) generation has advanced greatly, yet maintaining consistency remains challenging, especially for tuning-free long video generation. We attribute the consistency problem to cumulative deviations for long video generation at three levels: the random noise lacking correlation results initial deviation between frames; discrepancy in semantic feature tokens between denoising network blocks gradually accumulates as the frame count grows, leading to greater deviations; attention mechanisms struggle to capture global relationships across distant frames in long videos. To address these, we propose FreeMem, a tuning-free framework leveraging hierarchical memory update and injection: the noise memory stabilizes consistency by manipulating low and high frequency components in the initial noise space; the token memory combats inconsistency through adaptive fusion of historical and current semantic feature tokens between denoising network blocks; and the attention memory establishes persistent cache to model long-range relationships within self attention layers. Evaluated on VBench, FreeMem improves subject and background consistency matrics across various methods, offering a practical solution for low-cost, high-consistency long video generation.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 3cd6e0fc-9d83-4533-8e8c-140d07f2a26f

Builds on25

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines