Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset
Qingyan Bai, Qiuyu Wang, Hao Ouyang, Yue Yu, Hanlin Wang, Wen Wang, Ka Leong Cheng, Shuailei Ma, Yanhong Zeng, Zichen Liu, Yinghao Xu, Yujun Shen, Qifeng Chen
摘要
Instruction-based video editing promises to democratize content creation, yet its progress is severely hampered by the scarcity of large-scale, high-quality training data. We introduce Ditto, a holistic framework designed to tackle this fundamental challenge. At its heart, Ditto features a novel data generation pipeline that fuses the creative diversity of a leading image editor with an in-context video generator, overcoming the limited scope of existing models. To make this process viable, our framework resolves the prohibitive cost-quality trade-off by employing an efficient, distilled model architecture augmented by a temporal enhancer, which simultaneously reduces computational overhead and improves temporal coherence. Finally, to achieve full scalability, this entire pipeline is driven by an intelligent agent that crafts diverse instructions and rigorously filters the output, ensuring quality control at scale. Using this framework, we invested over 12,000 GPU-days to build Ditto-1M, a new dataset of one million high-fidelity video editing examples. We trained our model, Editto, on Ditto-1M with a curriculum learning strategy. The results demonstrate superior instruction-following ability and establish a new state-of-the-art in instruction-based video editing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- EditVerse: Unifying Image and Video Editing and Generation with In-Context LearningXuan Ju, Tianyu Wang, Yuqian Zhou, He Zhang 等ICLR 2026 · 被引用 56 次
- IVEBench: Modern Benchmark Suite for Instruction-Guided Video Editing AssessmentYinan Chen, Jiangning Zhang, Teng Hu, Yuxiang Zeng 等ICLR 2026 · 被引用 29 次
- VIVA: VLM-Guided Instruction-Based Video Editing with Reward OptimizationXiaoyan Cong, Haotian Yang, Angtian Wang, Yizhi Wang 等CVPR 2026 · 被引用 16 次
- EasyV2V: A High-quality Instruction-based Video Editing FrameworkJinjie Mai, Chaoyang Wang, Gordon Guocheng Qian, Willi Menapace 等CVPR 2026 · 被引用 12 次
- CoT-Edit: Let CoT Guide Instruction Video EditingSen Liang, Fengbin Guan, Youliang Zhang, Xin Li 等CVPR 2026 · 被引用 5 次
它引用的顶会 Paper38
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
相关 Paper
- Image Editing As Programs with Diffusion ModelsYujia Hu, Songhua Liu, Zhenxiong Tan, Xingyi Yang 等NeurIPS 2025 · 被引用 10 次
- VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded GenerationShoubin Yu, Difan Liu, Ziqiao Ma, Yicong Hong 等ICCV 2025 · 被引用 2 次
- InsViE-1M: Effective Instruction-Based Video Editing with Elaborate Dataset ConstructionYuhui Wu, Liyi Chen, Ruibin Li, Shihao Wang 等ICCV 2025 · 被引用 6 次
- GenCompositor: Generative Video Compositing with Diffusion TransformerShuzhou Yang, Xiaoyu Li, Xiaodong Cun, Guangzhi Wang 等ICLR 2026 · 被引用 10 次
- In-Context Generation with Regional Constraints for Instructional Video EditingZhongwei Zhang, Fuchen Long, Wei Li, Zhaofan Qiu 等ICML 2026
