VideoTitans: Scalable Video Prediction with Integrated Short- and Long-term Memory
Young-Jae Park, Minseok Seo, Hae-Gon Jeon
Abstract
Accurate video forecasting enables autonomous vehicles to anticipate hazards, robotics and surveillance systems to predict human intent, and environmental models to issue timely warnings for extreme weather events. However, existing methods remain limited: transformers rely on global attention with quadratic complexity, making them impractical for high-resolution, long-horizon video prediction, while convolutional and recurrent networks suffer from short-range receptive fields and vanishing gradients, losing key information over extended sequences. To overcome these challenges, we introduce VideoTitans , the first architecture to adapt the gradient-driven Titans memory—originally designed for language modelling to video prediction. VideoTitans integrates three core ideas: (i) a sliding-window attention core that scales linearly with sequence length and spatial resolution, (ii) an episodic memory that dynamically retains only informative tokens based on a gradient-based surprise signal, and (iii) a small set of persistent tokens encoding task-specific priors that stabilize training and enhance generalization. Extensive experiments on Moving-MNIST, Human3.6M, TrafficBJ and WeatherBench benchmarks show that VideoTitans consistently reduces computation (FLOPs) and achieves competitive visual fidelity compared to state-of-the-art recurrent, convolutional, and efficient-transformer methods. Comprehensive ablations confirm that each proposed component contributes significantly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7c81661-856e-4dca-898d-f212eeb668d5Cited by top-tier papers1
Ask how each one uses itBuilds on23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen et al.ICLR 2021 · 881 citations
- Synthesizer: Rethinking Self-Attention for Transformer ModelsYi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan et al.ICML 2021 · 399 citations
- Titans: Learning to Memorize at Test TimeAli Behrouz, Peilin Zhong, Vahab MirrokniNeurIPS 2025 · 368 citations
Related papers
- Memory Consolidation Enables Long-Context Video UnderstandingIvana Balazevic, Yuge Shi, Pinelopi Papalampidi, Rahma Chaabouni et al.ICML 2024 · 52 citations
- MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video RecognitionChao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan et al.CVPR 2022 · 158 citations
- Anticipative Video TransformerRohit Girdhar, Kristen GraumanICCV 2021 · 270 citations
- ATLAS: Learning to Optimally Memorize the Context at Test TimeAli Behrouz, Zeman Li, Praneeth Kacham, Majid Daliri et al.ICML 2026 · 57 citations
- CMC: Video Transformer Acceleration via CODEC Assisted Matrix CondensingZhuoran Song, Chunyu Qi, Fangxin Liu, Naifeng Jing et al.ASPLOS 2024 · 8 citations
