Lune

CVPR2024Top-tier venue

VideoBooth: Diffusion-based Video Generation with Image Prompts

Yuming Jiang, Tianxing Wu, Shuai Yang, Chenyang Si, Dahua Lin, Yu Qiao, Chen Change Loy, Ziwei Liu

2024Year
58Top-tier citations

Abstract

https://vchitect.github.io/VideoBooth-project/ Image Prompt Portrait of a dog, looks out the car window. Image Prompt Cat is looking at a laptop. A horse eating grass. Image Prompt Image Prompt Dog walking in the green farm 4k Elephant walk in the yellow grass of savannah Elephant drinking water in masai mara reserve, kenya Close up of cat on top of a vintage chair Horse grazes in snowy meadow * This work was done when Shuai Yang was in S-Lab, NTU. image prompts are fed into different cross-frame attention layers as additional keys and values. This extra spatial information refines the details in the first frame and then it is propagated to the remaining frames, which maintains temporal consistency. Extensive experiments demonstrate that VideoBooth achieves state-of-the-art performance in generating customized high-quality videos with subjects specified in image prompts. Notably, VideoBooth is a generalizable framework where a single model works for a wide range of image prompts with only feed-forward passes.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 7ac110f7-8919-4450-b603-38b8b8b291a2

Cited by top-tier papers58

Ask how each one uses it

Builds on48

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines