PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation
Qiyao Xue, Xiangyu Yin, Boyuan Yang, Wei Gao
Abstract
Multiple apples, bouncing Single apple, no bouncing Drawing content disappears Drawing with causality No water splashing Water splashing Calm water Flooding river No tumbling rock Rock tumbling No tea filling and steam of hot tea Tea is filling the cup with steam Figure 1. Left: videos generated by the current text-to-video generation model (CogVideoX-5B [50] ) cannot adhere to the real-world physical rules (described in brackets following the user prompt). Right: our method PhyT2V, when applied to the same model, better reflects the real-world physical knowledge.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5aad9b5b-020c-496e-8f7b-dec65e495a4eCited by top-tier papers21
- VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation ModelsXiangdong Zhang, Jiaqi Liao, Shaofeng Zhang, Fanqing Meng et al.NeurIPS 2025 · 98 citations
- WISA: World simulator assistant for physics-aware text-to-video generationJing Wang, Ao Ma, Ke Cao, Jun Zheng et al.NeurIPS 2025 · 93 citations
- Uni-CoT: Towards Unified Chain-of-Thought Reasoning Across Text and VisionLuozheng Qin, Jia Gong, Yuqing Sun, Tianjiao Li et al.ICLR 2026 · 55 citations
- NewtonGen: Physics-consistent and Controllable Text-to-Video Generation via Neural Newtonian DynamicsYu Yuan, Xijun Wang, Tharindu Wickremasinghe, Zeeshan Nadir et al.ICLR 2026 · 46 citations
- Inference-time Physics Alignment of Video Generative Models with Latent World ModelsJianhao Yuan, Xiaofeng Zhang, Felix Friedrich, Nicolas Beltran-Velez et al.CVPR 2026 · 32 citations
Builds on18
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li et al.ICLR 2024 · 467 citations
- Make-A-Video: Text-to-Video Generation without Text-Video DataUriel Singer, Adam Polyak, Thomas Hayes, Xi Yin et al.ICLR 2023 · 313 citations
- Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMsLing Yang, Zhaochen Yu, Chenlin Meng, Minkai Xu et al.ICML 2024 · 231 citations
Related papers
- VideoPhy: Evaluating Physical Commonsense for Video GenerationHritik Bansal, Zongyu Lin, Tianyi Xie, Zeshun Zong et al.ICLR 2025 · 1 citation
- Tora: Trajectory-oriented Diffusion Transformer for Video GenerationZhenghao Zhang, Junchao Liao, Menghao Li, Zuozhuo Dai et al.CVPR 2025
- PhysGen3D: Crafting a Miniature Interactive World from a Single ImageBoyuan Chen, Hanxiao Jiang, Shaowei Liu, Saurabh Gupta et al.CVPR 2025
- PhyWorldBench: A Comprehensive Evaluation of Physical Realism in Text-to-Video ModelsJing Gu, Xian Liu, Yu Zeng, Ashwin Nagarajan et al.ICLR 2026 · 29 citations
- Chain of Event-Centric Causal Thought for Physically Plausible Video GenerationZixuan Wang, Yixin Hu, Haolan Wang, Feng Chen et al.CVPR 2026 · 8 citations
