EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
Zun Wang, Jaemin Cho, Jialu Li, Han Lin, Jaehong Yoon, Yue Zhang, Mohit Bansal
摘要
Recent approaches for video generation with camera control often create anchor videos (i.e., rendered videos that approximate desired camera motions) to guide diffusion models as a structured prior, by rendering from estimated point clouds following camera trajectories. However, errors in point cloud and camera trajectory estimation often lead to inaccurate anchor videos with higher training cost and low efficiency, as the model is forced to compensate for rendering misalignments. To address these limitations, we introduce EPiC, an efficient and precise camera control learning framework that constructs well-aligned training anchor videos without the need for camera pose or point cloud estimation. Concretely, we create highly precise anchor videos by masking source videos based on first-frame visibility, which ensures strong alignment, eliminates the need for camera/point cloud estimation, and thus can be readily applied to any in-the-wild video. Furthermore, we introduce Anchor-ControlNet, a lightweight module that integrates anchor video guidance in visible regions to pretrained video diffusion models, with less than 1% of additional parameters. EPiC achieves efficient training with substantially fewer parameters, training steps, and less data, and generalizes robustly to anchor videos made with point clouds at test time, enabling precise 3D-informed camera control. EPiC achieves SoTA performance on RealEstate10K and MiraData for I2V camera control task. Notably, EPiC also exhibits strong zero-shot generalization to video-to-video (V2V) scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric ControlSixiao Zheng, Minghao Yin, Wenbo Hu, Xiaoyu Li 等CVPR 2026 · 被引用 27 次
- SpaceTimePilot: Generative Rendering of Dynamic Scenes Across Space and TimeZhening Huang, Hyeonho Jeong, Xuelin Chen, Yulia Gryaditskaya 等CVPR 2026 · 被引用 10 次
- EgoX: Egocentric Video Generation from a Single Exocentric VideoTaewoong Kang, Kinam Kim, Dohyeon Kim, Minho Park 等CVPR 2026 · 被引用 9 次
- LAMP: Language-Assisted Motion Planning for Controllable Video GenerationMuhammed Burak Kizil, Enes Şanlı, Niloy J. Mitra, Erkut Erdem 等CVPR 2026 · 被引用 4 次
它引用的顶会 Paper31
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video GeneratorsLevon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel 等ICCV 2023 · 被引用 800 次
- CAT3D: Create Anything in 3D with Multi-View Diffusion ModelsRuiqi Gao, Aleksander Holynski, Philipp Henzler, Arthur Brussee 等NeurIPS 2024 · 被引用 490 次
相关 Paper
- Latent-Reframe: Enabling Camera Control for Video Diffusion Models Without TrainingZhenghong Zhou, Jie An, Jiebo LuoICCV 2025 · 被引用 14 次
- Training-free Camera Control for Video GenerationChen Hou, Zhibo ChenICLR 2025
- VD3D: Taming Large Video Diffusion Transformers for 3D Camera ControlSherwin Bahmani, Ivan Skorokhodov, Aliaksandr Siarohin, Willi Menapace 等ICLR 2025
- ReRoPE: Repurposing RoPE for Relative Camera ControlChunyang Li, Yuanbo Yang, Jiahao Shao, Hongyu Zhou 等SIGGRAPH 2026 · 被引用 2 次
- ActCam: Zero-Shot Joint Camera and 3D Motion Control for Video GenerationOmar El Khalifi, Thomas Rossi, Oscar Fossey, Thibault Fouque 等SIGGRAPH 2026
