Taming Video Models for 3D and 4D Generation via Zero-Shot Camera Control
Chenxi Song, Yanming Yang, Tong Zhao, Ruibo Li, Chi Zhang
摘要
Video diffusion models have rich world priors, but their use in spatial tasks is limited by poor control, spatial-temporal inconsistent results, and entangled scene-camera dynamics. Current approaches, such as per-task fine-tuning or post-process warping, often introduce visual artifacts, fail to generalize, or incur high computational costs. We introduce WorldForge, a novel, training-free framework that operates purely at inference time to resolve these issues. Our method comprises three synergistic components. First, an intra-step refinement loop injects fine-grained motion guidance during the denoising process, iteratively correcting the output to ensure strict adherence to the target camera path. Second, an optical flow-based analysis identifies and isolates motion-related channels within the latent space. This allows our framework to selectively apply guidance, thereby decoupling motion from appearance and preserving visual fidelity. Third, a dual-path guidance strategy adaptively corrects for drift by comparing the guided generation against an unguided, reference denoising path, effectively neutralizing artifacts caused by misaligned structural inputs. Together, these components inject precise, trajectory-aligned control without model retraining, achieving accurate motion guidance and photorealistic synthesis. As a plug-and-play, model-agnostic solution, WorldForge demonstrates highly versatile generalizability. Beyond robust zero-shot 3D/4D generation, it readily empowers over a dozen diverse downstream applications, seamlessly enabling tasks like video editing, stabilization, and virtual try-on. Extensive experiments confirm state-of-the-art performance in trajectory adherence and perceptual quality, outperforming both training-dependent and inference-only baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- NeoVerse: Enhancing 4D World Model with in-the-wild Monocular VideosYuxue Yang, Lue Fan, Ziqi Shi, Junran Peng 等CVPR 2026 · 被引用 42 次
- FreeOrbit4D: Training-Free Arbitrary Camera Redirection for Monocular Videos via Foreground-Complete 4D ReconstructionWei Cao, Hao Zhang, Fengrui Tian, Yulun Wu 等SIGGRAPH 2026 · 被引用 4 次
- Free-Lunch Long Video Generation via Layer-Adaptive O.O.D CorrectionJiahao Tian, Chenxi Song, Wei Cheng, Chi ZhangCVPR 2026 · 被引用 3 次
它引用的顶会 Paper64
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
相关 Paper
- FlowDirector: Training-Free Flow Steering for Precise Text-to-Video EditingGuangzhao Li, Yanming Yang, Chenxi Song, Xiaohong Liu 等CVPR 2026 · 被引用 27 次
- Zo3T: Zero-Shot 3D-Aware Trajectory-Guided Image-to-Video Generation via Test-Time TrainingRuicheng Zhang, Jun Zhou, Zunnan Xu, Zihao Liu 等AAAI 2026 · 被引用 4 次
- Motion-Zero: A Zero-Shot Trajectory Control Framework of Moving Object for Diffusion-Based Video GenerationChanggu Chen, Junwei Shu, Gaoqi He, Changbo Wang 等AAAI 2025 · 被引用 1 次
- MotionFlow: Attention-Driven Motion Transfer in Video Diffusion ModelsTuna Han Salih Meral, Hidir Yesiltepe, Connor Dunlop, Pinar YanardagAAAI 2026
- DiffPerformer: Iterative Learning of Consistent Latent Guidance for Diffusion-Based Human Video GenerationChenyang Wang, Zerong Zheng, Tao Yu, Xiaoqian Lv 等CVPR 2024 · 被引用 3 次
