M3L: Language-based Video Editing via Multi-Modal Multi-Level Transformers
Tsu-Jui Fu, Xin Eric Wang, Scott T. Grafton, Miguel P. Eckstein, William Yang Wang
摘要
Video editing tools are widely used nowadays for digital design. Although the demand for these tools is high, the prior knowledge required makes it difficult for novices to get started. Systems that could follow natural language instructions to perform automatic editing would significantly improve accessibility. This paper introduces the language-based video editing (LBVE) task, which allows the model to edit, guided by text instruction, a source video into a target video. LBVE contains two features: 1) the scenario of the source video is preserved instead of generating a completely different video; 2) the semantic is presented differently in the target video, and all changes are controlled by the given instruction. We propose a Multi-Modal Multi-Level Transformer (M <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">3</sup> L) to carry out LBVE. M <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">3</sup> L dynamically learns the correspondence between video perception and language semantic at different levels, which benefits both the video understanding and video frame synthesis. We build three new datasets for evaluation, including two diagnostic and one from natural videos with human-labeled text. Extensive experimental results show that M <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">3</sup> L is effective for video editing and that LBVE can lead to a new field toward vision-and-language research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Shatter and Gather: Learning Referring Image Segmentation with Text SupervisionDongwon Kim, Namyup Kim, Cuiling Lan, Suha KwakICCV 2023 · 被引用 29 次
- Learning Cross-Modal Affinity for Referring Video Object Segmentation Targeting Limited SamplesGuanghui Li, Mingqi Gao, Heng Liu, Xiantong Zhen 等ICCV 2023 · 被引用 6 次
- Vulnerability-Aware Spatio-Temporal Learning for Generalizable Deepfake Video DetectionDat Nguyen, Marcella Astrid, Anis Kacem, Enjie Ghorbel 等ICCV 2025 · 被引用 5 次
- SynopGround: A Large-Scale Dataset for Multi-Paragraph Video Grounding from TV Dramas and SynopsesChaolei Tan, Zihang Lin, Junfu Pu, Zhongang Qi 等ACM MM 2024 · 被引用 2 次
- AltFreezing for More General Video Face Forgery DetectionZhendong Wang, Jianmin Bao, Wengang Zhou, Weilun Wang 等CVPR 2023
它引用的顶会 Paper5
- Scaling Autoregressive Video ModelsDirk Weissenborn, Oscar Täckström, Jakob UszkoreitICLR 2020 · 被引用 252 次
- Arbitrary Video Style Transfer via Multi-Channel CorrelationYingying Deng, Fan Tang, Weiming Dong, Haibin Huang 等AAAI 2021 · 被引用 197 次
- Tell, Draw, and Repeat: Generating and Modifying Images Based on Continual Linguistic InstructionAlaaeldin El-Nouby, Shikhar Sharma, Hannes Schulz, R. Devon Hjelm 等ICCV 2019 · 被引用 128 次
- SSCR: Iterative Language-Based Image Editing via Self-Supervised Counterfactual ReasoningTsu-Jui Fu, Xin Wang, Scott T. Grafton, Miguel P. Eckstein 等EMNLP 2020 · 被引用 26 次
- Disentangling Physical Dynamics From Unknown Factors for Unsupervised Video PredictionVincent Le Guen, Nicolas ThomeCVPR 2020
相关 Paper
- MiVE: Multiscale Vision-language features for reference-guided video EditingTong Wang, Meng Zou, WU CHENGJING, Xiaochao Qu 等ICML 2026
- Modeling Motion with Multi-Modal Features for Text-Based Video SegmentationWangbo Zhao, Kai Wang, Xiangxiang Chu, Fuzhao Xue 等CVPR 2022 · 被引用 23 次
- VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded GenerationShoubin Yu, Difan Liu, Ziqiao Ma, Yicong Hong 等ICCV 2025 · 被引用 2 次
- UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text EditingLichen Ma, Xiaolong Fu, Gaojing Zhou, Zipeng Guo 等AAAI 2026 · 被引用 1 次
- Guiding Instruction-based Image Editing via Multimodal Large Language ModelsTsu-Jui Fu, Wenze Hu, Xianzhi Du, William Yang Wang 等ICLR 2024 · 被引用 173 次
