Generative Video Motion Editing with 3D Point Tracks
Yao-Chih Lee, Zhoutong Zhang, Jiahui Huang, Jui-Hsien Wang, Joon-Young Lee, Jia-Bin Huang, Eli Shechtman, Zhengqi Li
摘要
Camera and object motions are central to a video's narrative. However, precisely editing these captured motions remains a significant challenge, especially under complex object movements. Current motion-controlled image-to-video (I2V) approaches often lack full-scene context for consistent video editing, while video-to-video (V2V) methods provide viewpoint changes or basic object translation, but offer limited control over fine-grained object motion. We present a track-conditioned V2V framework that enables joint editing of camera and object motion. We achieve this by conditioning a video generation model on a source video and paired 3D point tracks representing source and target motions. These 3D tracks establish sparse correspondences that transfer rich context from the source video to new motions while preserving spatiotemporal coherence. Crucially, compared to 2D tracks, 3D tracks provide explicit depth cues, allowing the model to resolve depth order and handle occlusions for precise motion editing. Trained in two stages on synthetic and real data, our model supports diverse motion edits, including joint camera/object manipulation, motion transfer, and non-rigid deformation, unlocking new creative potential in video editing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- 4RC: 4D Reconstruction via Conditional Querying Anytime and AnywhereYihang Luo, Shangchen Zhou, Yushi Lan, Xingang Pan 等ICML 2026 · 被引用 12 次
- FreeOrbit4D: Training-Free Arbitrary Camera Redirection for Monocular Videos via Foreground-Complete 4D ReconstructionWei Cao, Hao Zhang, Fengrui Tian, Yulun Wu 等SIGGRAPH 2026 · 被引用 4 次
- MotiMotion: Motion-Controlled Video Generation with Visual ReasoningHsin-Ying Lee, Hanwen Jiang, Yiqun Mei, Jing Shi 等ICML 2026
它引用的顶会 Paper79
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
相关 Paper
- MotionV2V: Editing Motion in a VideoRyan D. Burgert, Charles Herrmann, Forrester Cole, Michael S. Ryoo 等CVPR 2026 · 被引用 13 次
- Motion Prompting: Controlling Video Generation with Motion TrajectoriesDaniel Geng, Charles Herrmann, Junhwa Hur, Forrester Cole 等CVPR 2025
- UniScene-MoTion: Unified Scene & Motion-aware Diffusion Transition FrameworkRui Jiang, Chongmian Wang, Xinghe Fu, Yehao Lu 等AAAI 2026
- TrackGo: A Flexible and Efficient Method for Controllable Video GenerationHaitao Zhou, Chuang Wang, Rui Nie, Jinlin Liu 等AAAI 2025 · 被引用 41 次
- Time-to-Move: Training-Free Motion-Controlled Video Generation via Dual-Clock DenoisingAssaf Singer, Noam Rotstein, Amir Mann, Ron Kimmel 等ICLR 2026 · 被引用 13 次
