Lune

ACM MM2025顶会

Text Prompted Spatiotemporal Sequence Prediction with Text-Vision Prompt Refiner and Masked Diffusion Transformers

Yechao Xu, Zhengxing Sun, Qian Li, Yunhan Sun

2025年份

摘要

Classical spatiotemporal sequence prediction tasks are designed to forecast future image sequences based on historical observations. However, the inherent unpredictability of future events often renders this process uncontrollable due to infinite possibilities in nature, limiting broader applicability of this technology. In this study, we explore the utilization of text prompts to constrain probabilistic space of future outcomes, resulting more controllable future prediction complying with user intent. We primarily address two critical challenges in this research setting: (i) text-vision misalignment, where embeddings extracted by text pre-trained models are not strictly aligned with visual embeddings, leading to predictions semantically irrelevant to text prompts. (ii) Spatiotemporal modeling distortion, where the fixed observation interval during training causes the model to produce unrealistic results when reasoning longer time dimensions. To tackle these issues, we propose a text-prompted spatiotemporal sequence prediction (TPS2P) model, leveraging historical observations and textual prompts to predict probabilistic future outcomes. In this model, a text-vision prompt refiner (TV-Refiner) is introduced to provide aligned textual and historical visual embeddings for integrating the denoising diffusion prediction process. Additionally, a spatiotemporal-masked diffusion transformer (StMDiT) is proposed by exploiting masked attention in constituting spatial and temporal self-attention modules within latent diffusion processes, enabling the model to observe more sequences of varying spatiotemporal patterns during training. We conduct extensive experiments on Something-Something V2 (Sthv2) and BridgeData datasets. Reported results demonstrate that our TPS2P predicts more accurate and high-quality future sequences, more user-intent compliant by textual controllability.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖