Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control
Zhengfei Kuang, Shengqu Cai, Hao He, Yinghao Xu, Hongsheng Li, Leonidas J. Guibas, Gordon Wetzstein
Abstract
Research on video generation has recently made tremendous progress, enabling high-quality videos to be generated from text prompts or images. Adding control to the video generation process is an important goal moving forward and recent approaches that condition video generation models on camera trajectories make strides towards it. Yet, it remains challenging to generate a video of the same scene from multiple different camera trajectories. Solutions to this multi-video generation problem could enable large-scale 3D scene generation with editable camera trajectories, among other applications. We introduce collaborative video diffusion (CVD) as an important step towards this vision. The CVD framework includes a novel cross-video synchronization module that promotes consistency between corresponding frames of the same video rendered from different camera poses using an epipolar attention mechanism. Trained on top of a state-of-the-art camera-control module for video generation, CVD generates multiple videos rendered from different camera trajectories with significantly better consistency than baselines, as shown in extensive experiments. Project page: https://collaborativevideodiffusion.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c84b3ea0-1358-4acb-8584-dd98e1bf1fc7Cited by top-tier papers60
- SpatialVID: A Large-Scale Video Dataset with Spatial AnnotationsJiahao Wang, Yufeng Yuan, Rujie Zheng, Youtian Lin et al.CVPR 2026 · 72 citations
- Vivid-ZOO: Multi-View Video Generation with Diffusion ModelBing Li, Cheng Zheng, Wenxuan Zhu, Jinjie Mai et al.NeurIPS 2024 · 48 citations
- EasyCreator: Empowering 4D Creation through Video InpaintingYue Ma, Kunyu Feng, Xinhua Zhang, Hongyu Liu et al.ICLR 2026 · 47 citations
- Learning Video Generation for Robotic Manipulation with Collaborative Trajectory ControlXiao Fu, Xintao Wang, Xian Liu, Jianhong Bai et al.ICLR 2026 · 37 citations
- Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-DistillationSherwin Bahmani, Tianchang Shen, Jiawei Ren, Jiahui Huang et al.ICLR 2026 · 33 citations
Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- CameraSquad: Achieving Content Consistency in Parallel Multi-Trajectory Camera-Controlled Video GenerationZhufeng Xu, Xuan Gao, Bailin Deng, Yikang Ding et al.SIGGRAPH 2026
- AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion TransformersSherwin Bahmani, Ivan Skorokhodov, Guocheng Qian, Aliaksandr Siarohin et al.CVPR 2025
- Trajectory attention for fine-grained video motion controlZeqi Xiao, Wenqi Ouyang, Yifan Zhou, Shuai Yang et al.ICLR 2025
- CameraCtrl: Enabling Camera Control for Video Diffusion ModelsHao He, Yinghao Xu, Yuwei Guo, Gordon Wetzstein et al.ICLR 2025
- Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated AttentionDejia Xu, Yifan Jiang, Chen Huang, Liangchen Song et al.ICML 2025
