Performance over Random: A Robust Evaluation Protocol for Video Summarization Methods
Evlampios Apostolidis, Eleni Adamantidou, Alexandros I. Metsai, Vasileios Mezaris, Ioannis Patras
摘要
This paper proposes a new evaluation approach for video summarization algorithms. We start by studying the currently established evaluation protocol; this protocol, defined over the ground-truth annotations of the SumMe and TVSum datasets, quantifies the agreement between the user-defined and the automatically-created summaries with F-Score, and reports the average performance on a few different training/testing splits of the used dataset. We evaluate five publicly-available summarization algorithms under a large-scale experimental setting with 50 randomly-created data splits. We show that the results reported in the papers are not always congruent with their performance on the large-scale experiment, and that the F-Score cannot be used for comparing algorithms evaluated on different splits. We also show that the above shortcomings of the established evaluation protocol are due to the significantly varying levels of difficulty among the utilized splits, that affect the outcomes of the evaluations. Further analysis of these findings indicates a noticeable performance correlation among all algorithms and a random summarizer. To mitigate these shortcomings we propose an evaluation protocol that makes estimates about the difficulty of each used data split and utilizes this information during the evaluation process. Experiments involving different evaluation settings demonstrate the increased representativeness of performance results when using the proposed evaluation approach, and the increased reliability of comparisons when the examined methods have been evaluated on different data splits.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Video Summarization Using Denoising Diffusion Probabilistic ModelZirui Shang, Yubo Zhu, Hongxi Li, Shuo Yang 等AAAI 2025 · 被引用 4 次
- SummDiff: Generative Modeling of Video Summarization with DiffusionKwanseok Kim, Jaehoon Hahm, Sumin Kim, Jinhwan Sul 等ICCV 2025 · 被引用 1 次
- Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human EvaluationYixin Liu, Alexander R. Fabbri, Pengfei Liu, Yilun Zhao 等ACL 2023 · 被引用 50 次
- SEM-F1: an Automatic Way for Semantic Evaluation of Multi-Narrative Overlap Summaries at ScaleNaman Bansal, Mousumi Akter, Shubhra Kanti Karmaker SantuEMNLP 2022 · 被引用 2 次
- STREAM: Spatio-TempoRal Evaluation and Analysis Metric for Video Generative ModelsPum Jun Kim, Seojun Kim, Jaejun YooICLR 2024 · 被引用 11 次
