Latent Knowledge-Guided Video Diffusion for Scientific Phenomena Generation from a Single Initial Frame
Qinglong Cao, Xirui Li, Ding Wang, Chao Ma, Yuntian Chen, Xiaokang Yang
摘要
Video diffusion models have achieved impressive results in natural scene generation, yet they struggle to generalize to scientific phenomena such as fluid simulations and meteorological processes, where underlying dynamics are governed by scientific laws. These tasks pose unique challenges, including severe domain gaps, limited training data, and the lack of descriptive language annotations. To handle this dilemma, we extracted the latent scientific phenomena knowledge and further proposed a fresh framework that teaches video diffusion models to generate scientific phenomena from a single initial frame. Particularly, static knowledge is extracted via pre-trained masked autoencoders, while dynamic knowledge is derived from pre-trained optical flow prediction. Subsequently, based on the aligned spatial relations between the CLIP vision and language encoders, the visual embeddings of scientific phenomena, guided by latent scientific phenomena knowledge, are projected to generate the pseudo-language prompt embeddings in both spatial and frequency domains. By incorporating these prompts and fine-tuning the video diffusion model, we enable the generation of videos that better adhere to scientific laws. Extensive experiments on both computational fluid dynamics simulations and real-world typhoon observations demonstrate the effectiveness of our approach, achieving superior fidelity and consistency across diverse scientific scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Inference-time Physics Alignment of Video Generative Models with Latent World ModelsJianhao Yuan, Xiaofeng Zhang, Felix Friedrich, Nicolas Beltran-Velez 等CVPR 2026 · 被引用 32 次
- SeeU: Seeing the Unseen World via 4D Dynamics-aware GenerationYu Yuan, Tharindu Wickremasinghe, Zeeshan Nadir, Xijun Wang 等CVPR 2026 · 被引用 3 次
- Pixel2Phys: Distilling Governing Laws from Visual DynamicsRuikun Li, Jun Yao, Yingfan Hua, Shixiang Tang 等CVPR 2026 · 被引用 2 次
- Omni-Weather: A Unified Multimodal Model for Weather Radar Understanding and GenerationZhiwang Zhou, Yuandong Pu, Xuming He, Yidi Liu 等ICLR 2026
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- MoAlign: Motion-Centric Representation Alignment for Video Diffusion ModelsAritra Bhowmik, Denis Korzhenkov, Cees G. M. Snoek, Amir Habibian 等ICLR 2026 · 被引用 15 次
- VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical PriorXindi Yang, Baolu Li, Yiming Zhang, Zhenfei Yin 等ICCV 2025 · 被引用 8 次
- Bringing Real-World Relations into Video Generation with Graph-Structured KnowledgeJoonhyung Park, Jaeyun Song, Sihwan Park, Eunho YangACL 2026
- MeDM: Mediating Image Diffusion Models for Video-to-Video Translation with Temporal Correspondence GuidanceErnie Chu, Tzuhsuan Huang, Shuo-Yen Lin, Jun-Cheng ChenAAAI 2024 · 被引用 25 次
- Chain of Event-Centric Causal Thought for Physically Plausible Video GenerationZixuan Wang, Yixin Hu, Haolan Wang, Feng Chen 等CVPR 2026 · 被引用 8 次
