PoseTraj: Pose-Aware Trajectory Control in Video Diffusion
Longbin Ji, Lei Zhong, Pengfei Wei, Changjian Li
摘要
Recent advancements in trajectory-guided video generation have achieved notable progress. However, existing models still face challenges in generating object motions with potentially changing 6D poses under wide-range rotations, due to limited 3D understanding. To address this problem, we introduce PoseTraj, a pose-aware video dragging model for generating 3D-aligned motion from 2D trajectories. Our method adopts a novel two-stage pose-aware pretraining framework, improving 3D understanding across diverse trajectories. Specifically, we propose a large-scale synthetic dataset PoseTraj-10K, containing 10k videos of objects following rotational trajectories, and enhance the model perception of object pose changes by incorporating 3D bounding boxes as intermediate supervision signals. Following this, we fine-tune the trajectory-controlling module on real-world videos, applying an additional camera-disentanglement module to further refine motion accuracy. Experiments on various benchmark datasets demonstrate that our method not only excels in 3D pose-aligned dragging for rotational trajectories but also outperforms existing baselines in trajectory accuracy and video quality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Captain Safari: A World Engine with Pose-Aligned 3D MemoryYu-Cheng Chou, Xingrui Wang, Yitong Li, Jiahao Wang 等CVPR 2026
- Go-with-the-Track: Video Compositing and Motion Control with Point TrackingKoichi Namekata, Yash Kant, Zhizheng Liu, Ryan D. Burgert 等SIGGRAPH 2026
它引用的顶会 Paper22
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Structure and Content-Guided Video Synthesis with Diffusion ModelsPatrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog 等ICCV 2023 · 被引用 733 次
- Follow Your Pose: Pose-Guided Text-to-Video Generation Using Pose-Free VideosYue Ma, Yingqing He, Xiaodong Cun, Xintao Wang 等AAAI 2024 · 被引用 318 次
相关 Paper
- SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video GenerationGuiyu Zhang, Yabo Chen, Xunzhi Xiang, Junchao Huang 等CVPR 2026 · 被引用 8 次
- Towards Physical Understanding in Video Generation: A 3D Point Regularization ApproachYunuo Chen, Junli Cao, Vidit Goel, Sergei Korolev 等NeurIPS 2025 · 被引用 9 次
- PoseAnything: General Pose-guided Video Generation with Part-aware Temporal CoherenceRuiyan Wang, Teng Hu, Kaihui Huang, Zihan Su 等CVPR 2026
- 3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video GenerationXiao Fu, Xian Liu, Xintao Wang, Sida Peng 等ICLR 2025
- MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory GuidanceQuanhao Li, Zhen Xing, Rui Wang, Hui Zhang 等ICCV 2025 · 被引用 10 次
