RACCooN: Versatile Instructional Video Editing with Auto-Generated Narratives
Jaehong Yoon, Shoubin Yu, Mohit Bansal
Abstract
Recent video generative models primarily rely on detailed, labor-intensive text prompts for tasks, like inpainting or style editing, limiting adaptability for personal/raw videos. This paper proposes RACCOON, a versatile and userfriendly video-to-paragraph-to-video editing method, supporting diverse video editing capabilities, such as removal, addition, and modification, through a unified pipeline. RAC-COON consists of two main stages: Video-to-Paragraph (V2P), which automatically generates structured descriptions of scene and object details, and Paragraph-to-Video (P2V), where users can refine these to guide a video diffusion model for flexible content edits, including removing, changing, or adding objects. Key contributions of RACCOON include: (1) A multi-granular spatiotemporal pooling strategy for structured video understanding, capturing both global context and fine-grained object details to enable precise text-based video editing without complex human annotations. (2) A video generative model fine-tuned on a curated video-paragraph-mask dataset for improved editing and inpainting. (3) The ability to generate new objects by forecasting motion via auto-generated mask planning. In the end, users can easily edit complex videos with RAC-CooN's automatic explanations and guidance. We demonstrate its versatile capabilities in video-to-paragraph generation (up to 9.4%p Ò improvement in human evaluations), video content editing (relative 49.7% Ó in FVD).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on23
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- VideoComposer: Compositional Video Synthesis with Motion ControllabilityXiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen et al.NeurIPS 2023 · 579 citations
Related papers
- VideoCoF: Unified Video Editing with Temporal ReasonerXiangpeng Yang, Ji Xie, Yiyuan Yang, Yue Ma et al.CVPR 2026 · 6 citations
- Video-P2P: Video Editing with Cross-Attention ControlShaoteng Liu, Yuechen Zhang, Wenbo Li, Zhe Lin et al.CVPR 2024 · 99 citations
- VideoDirector: Precise Video Editing via Text-to-Video ModelsYukun Wang, Longguang Wang, Zhiyuan Ma, Qibin Hu et al.CVPR 2025
- Structure and Content-Guided Video Synthesis with Diffusion ModelsPatrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog et al.ICCV 2023 · 733 citations
- CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and CompatibilityBojia Zi, Shihao Zhao, Xianbiao Qi, Jianan Wang et al.AAAI 2025 · 6 citations
