VideoDiff: Human-AI Video Co-Creation with Alternatives
Mina Huh, Ding Li, Kim Pimmel, Hijung Valentina Shin, Amy Pavel, Mira Dontcheva
Abstract
To make an engaging video, people sequence interesting moments and add visuals such as B-rolls or text. While video editing requires time and effort, AI has recently shown strong potential to make editing easier through suggestions and automation. A key strength of generative models is their ability to quickly generate multiple variations, but when provided with many alternatives, creators struggle to compare them to find the best fit. We propose VideoDiff, an AI video editing tool designed for editing with alternatives. With VideoDiff, creators can generate and review multiple AI recommendations for each editing process: creating a rough cut, inserting B-rolls, and adding text effects. VideoDiff simplifies comparisons by aligning videos and highlighting differences through timelines, transcripts, and video previews. Creators have the flexibility to regenerate and refine AI suggestions as they compare alternatives. Our study participants (N=12) could easily compare and customize alternatives, creating more satisfying results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e96a8d0a-0a3e-4684-9117-a670f7eb675eCited by top-tier papers4
- GenTune: Toward Traceable Prompts to Improve Controllability of Image Refinement in Environment DesignWen-Fan Wang, Ting-Ying Lee, Chien-Ting Lu, Che-Wei Hsu et al.UIST 2025 · 4 citations
- VidTune: Creating Video Soundtracks with Generative Music and Video-Based ThumbnailsMina Huh, C. Ailie Fraser, Dingzeyu Li, Mira Dontcheva et al.CHI 2026 · 1 citation
- Designing Multi-Robot Ground Video Sensemaking with Public Safety ProfessionalsPuqi Zhou, Ali Asgarov, Aafiya Hussain, Wonjoon Park et al.CHI 2026 · 1 citation
- Vidmento: Creating Video Stories through Context-Aware Expansion with Generative VideoCatherine Yeh, Anh Truong, Mira Dontcheva, Bryan WangCHI 2026 · 1 citation
Builds on33
- AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model PromptsTongshuang Wu, Michael Terry, Carrie Jun CaiCHI 2022 · 465 citations
- Creating Augmented and Virtual Reality Applications: Current Practices, Challenges, and OpportunitiesNarges Ashtari, Andrea Bunt, Joanna McGrenere, Michael Nebeling et al.CHI 2020 · 274 citations
- Co-Writing Screenplays and Theatre Scripts with Language Models: Evaluation by Industry ProfessionalsPiotr Mirowski, Kory W. Mathewson, Jaylen Pittman, Richard EvansCHI 2023 · 235 citations
- The Effects of Generative AI on Design Fixation and Divergent ThinkingSamangi Wadinambiarachchi, Ryan M. Kelly, Saumya Pareek, Qiushi Zhou et al.CHI 2024 · 186 citations
- Promptify: Text-to-Image Generation through Interactive Prompt Exploration with Large Language ModelsStephen Brade, Bryan Wang, Maurício Sousa, Sageev Oore et al.UIST 2023 · 179 citations
Related papers
- SoundStager: Interactive Design of Story-Driven GenAI Soundscapes for VideoSuhyeon Yoo, Adolfo Hernandez Santisteban, Prem Seetharaman, Justin Salamon et al.CHI 2026 · 2 citations
- TokenFlow: Consistent Diffusion Features for Consistent Video EditingMichal Geyer, Omer Bar-Tal, Shai Bagon, Tali DekelICLR 2024 · 439 citations
- Structure and Content-Guided Video Synthesis with Diffusion ModelsPatrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog et al.ICCV 2023 · 733 citations
- EditBoard: Towards a Comprehensive Evaluation Benchmark for Text-Based Video Editing ModelsYupeng Chen, Penglin Chen, Xiaoyu Zhang, Yixian Huang et al.AAAI 2025 · 5 citations
- FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video SynthesisFeng Liang, Bichen Wu, Jialiang Wang, Licheng Yu et al.CVPR 2024
