Towards Physical Understanding in Video Generation: A 3D Point Regularization Approach
Yunuo Chen, Junli Cao, Vidit Goel, Sergei Korolev, Chenfanfu Jiang, Jian Ren, Sergey Tulyakov, Anil Kag
Abstract
We present a novel video generation framework that integrates 3-dimensional geometry and dynamic awareness. To achieve this, we augment 2D videos with 3D point trajectories and align them in pixel space. The resulting 3D-aware video dataset, PointVid, is then used to fine-tune a latent diffusion model, enabling it to track 2D objects with 3D Cartesian coordinates. Building on this, we regularize the shape and motion of objects in the video to eliminate undesired artifacts, e.g., non-physical deformation. Consequently, we enhance the quality of generated RGB videos and alleviate common issues like object morphing, which are prevalent in current video models due to a lack of shape awareness. With our 3D augmentation and regularization, our model is capable of handling contact-rich scenarios such as task-oriented videos, where 3D information is essential for perceiving shape and motion of interacting solids. Our method can be seamlessly integrated into existing video diffusion models to improve their visual plausibility.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5353087d-d5e1-4e57-8b6a-b0ac6003fe0cCited by top-tier papers4
- VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation ModelsXiangdong Zhang, Jiaqi Liao, Shaofeng Zhang, Fanqing Meng et al.NeurIPS 2025 · 98 citations
- RoboScape: Physics-informed Embodied World ModelYu Shang, Xin Zhang, Yinzhou Tang, Lei Jin et al.NeurIPS 2025 · 43 citations
- MoAlign: Motion-Centric Representation Alignment for Video Diffusion ModelsAritra Bhowmik, Denis Korzhenkov, Cees G. M. Snoek, Amir Habibian et al.ICLR 2026 · 15 citations
- UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video GenerationJiehui Huang, Yuechen Zhang, Xu He, Yuan Gao et al.CVPR 2026 · 12 citations
Builds on22
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- VideoPoet: A Large Language Model for Zero-Shot Video GenerationDan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama et al.ICML 2024 · 464 citations
- Make-A-Video: Text-to-Video Generation without Text-Video DataUriel Singer, Adam Polyak, Thomas Hayes, Xi Yin et al.ICLR 2023 · 313 citations
Related papers
- GeoVideo: Introducing Geometric Regularization into Video Generation ModelYunpeng Bai, Shaoheng Fang, Chaohui Yu, Fan Wang et al.NeurIPS 2025 · 18 citations
- Autodecoding Latent 3D Diffusion ModelsEvangelos Ntavelis, Aliaksandr Siarohin, Kyle Olszewski, Chaoyang Wang et al.NeurIPS 2023 · 65 citations
- Vivid-ZOO: Multi-View Video Generation with Diffusion ModelBing Li, Cheng Zheng, Wenxuan Zhu, Jinjie Mai et al.NeurIPS 2024 · 48 citations
- ManipDreamer3D: Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D TrajectoryYing Li, Xiaobao Wei, Xiaowei Chi, Yuming Li et al.AAAI 2026
- Multi-Identity Human Image Animation with Structural Video DiffusionZhenzhi Wang, Yixuan Li, Yanhong Zeng, Yuwei Guo et al.ICCV 2025 · 1 citation
