VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment
Shangkun Sun, Xiaoyu Liang, Songlin Fan, Wenxu Gao, Wei Gao
摘要
Text-driven video editing has recently experienced rapid development. Despite this, evaluating edited videos remains a considerable challenge. Current metrics tend to fail to align with human perceptions, and effective quantitative metrics for video editing are still notably absent. To address this, we introduce VE-Bench, a benchmark suite tailored to the assessment of text-driven video editing. This suite includes VE-Bench DB, a video quality assessment (VQA) database for video editing. VE-Bench DB encompasses a diverse set of source videos featuring various motions and subjects, along with multiple distinct editing prompts, editing results from 8 different models, and the corresponding Mean Opinion Scores (MOS) from 24 human annotators. Based on VE-Bench DB, we further propose VE-Bench QA, a quantitative human-aligned measurement for the text-driven video editing task. In addition to the aesthetic, distortion, and other visual quality indicators that traditional VQA methods emphasize, VE-Bench QA focuses on the text-video alignment and the relevance modeling between source and edited videos. It introduces a new assessment network for video editing that attains superior performance in alignment with human preferences.To the best of our knowledge, VE-Bench introduces the first quality assessment dataset for video editing and proposes an effective subjective-aligned quantitative metric for this domain. All models, data, and code will be publicly available to the community.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- IVEBench: Modern Benchmark Suite for Instruction-Guided Video Editing AssessmentYinan Chen, Jiangning Zhang, Teng Hu, Yuxiang Zeng 等ICLR 2026 · 被引用 29 次
- PPLLaVA: Varied Video Sequence Understanding With Prompt GuidanceShangkun Sun, Ruyang Liu, Haoran Tang, Yixiao Ge 等ICLR 2026 · 被引用 21 次
- AutoDrive-P3: Unified Chain of Perception-Prediction-Planning Thought via Reinforcement Fine-TuningYuqi Ye, Zijian Zhang, Junhong Lin, Shangkun Sun 等ICLR 2026 · 被引用 16 次
- Flow4Agent: Long-form Video Understanding via Motion Prior from Optical FlowRuyang Liu, Shangkun Sun, Haoran Tang, Wei Gao 等ICCV 2025 · 被引用 15 次
- UniVBench: Towards Unified Evaluation for Video Foundation ModelsJianhui Wei, Xiaotian Zhang, Yichen Li, Yuan Wang 等CVPR 2026 · 被引用 12 次
它引用的顶会 Paper33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 被引用 1,550 次
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong 等NeurIPS 2023 · 被引用 1,310 次
相关 Paper
- VBench: Comprehensive Benchmark Suite for Video Generative ModelsZiqi Huang, Yinan He, Jiashuo Yu, Fan Zhang 等CVPR 2024
- VMBench: A Benchmark for Perception-Aligned Video Motion GenerationXinran Ling, Chen Zhu, Meiqi Wu, Hangyu Li 等ICCV 2025 · 被引用 2 次
- Subjective-Aligned Dataset and Metric for Text-to-Video Quality AssessmentTengchuan Kou, Xiaohong Liu, Zicheng Zhang, Chunyi Li 等ACM MM 2024 · 被引用 29 次
- VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality EvaluationLongteng Jiang, Dandan Zheng, Qianqian Qiao, Heng Huang 等CVPR 2026 · 被引用 2 次
- LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMsZitong Xu, Huiyu Duan, Bingnan Liu, Guangji Ma 等ACM MM 2025 · 被引用 2 次
