Lune

EMNLP2025顶会

RACCooN: Versatile Instructional Video Editing with Auto-Generated Narratives

Jaehong Yoon, Shoubin Yu, Mohit Bansal

2025年份
2被引次数
1顶会引用

摘要

Recent video generative models primarily rely on detailed, labor-intensive text prompts for tasks, like inpainting or style editing, limiting adaptability for personal/raw videos. This paper proposes RACCOON, a versatile and userfriendly video-to-paragraph-to-video editing method, supporting diverse video editing capabilities, such as removal, addition, and modification, through a unified pipeline. RAC-COON consists of two main stages: Video-to-Paragraph (V2P), which automatically generates structured descriptions of scene and object details, and Paragraph-to-Video (P2V), where users can refine these to guide a video diffusion model for flexible content edits, including removing, changing, or adding objects. Key contributions of RACCOON include: (1) A multi-granular spatiotemporal pooling strategy for structured video understanding, capturing both global context and fine-grained object details to enable precise text-based video editing without complex human annotations. (2) A video generative model fine-tuned on a curated video-paragraph-mask dataset for improved editing and inpainting. (3) The ability to generate new objects by forecasting motion via auto-generated mask planning. In the end, users can easily edit complex videos with RAC-CooN's automatic explanations and guidance. We demonstrate its versatile capabilities in video-to-paragraph generation (up to 9.4%p Ò improvement in human evaluations), video content editing (relative 49.7% Ó in FVD).

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper23

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖