FPS-Bench: A Benchmark for High Frame-Rate Video Understanding
Rohan Choudhury, Jean Sebastien Dandurand, Kai Qiu, Kshitij Madhav Bhat, Kartik Sharma, Liza Dahiya, Yizhou Zhao, Souraja Kundu, Chun-Hsien Lin, Kris M. Kitani, László A. Jeni
Abstract
Modern video-language models are typically trained on videos downsampled to low frames-per-second (FPS), and the most commonly used evaluation benchmarks are designed for low-FPS input as well. To address this shortcoming, we present FPS-Bench, a large video question-answering benchmark designed to evaluate VLMs’ capabilities to understand video at high-frame rates. We introduce a new metric, the minimum frames-per-second (minFPS), which measures the minimum frame-rate required to solve a given question. While existing benchmarks require <1 minFPS, we rigorously curate more than 1000 questions from a diverse source of videos and manually verify minFPS for each example, leading to a benchmark that requires watching videos at on average 7 FPS to solve. Our evaluation of several state-of-the-art VLMs shows that they are severely lacking, achieving QA accuracy of 30% in the FPS-Bench multiple-choice task, while humans achieve 72% accuracy. We believe that FPS-Bench will serve as a valuable tool for improving frontier-level VLMs and will release all data and code.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on19
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Just Ask: Learning to Answer Questions from Millions of Narrated VideosAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev et al.ICCV 2021 · 345 citations
Related papers
- HERBench: A Benchmark for Multi-Evidence Integration in Video Question AnsweringDan Ben Ami, Gabriele Serussi, Kobi Cohen, Chaim BaskinCVPR 2026 · 3 citations
- V2P-Bench: Evaluating Video-Language Understanding with Visual Prompts for Better Human-Model InteractionYiming Zhao, Yu Zeng, Yukun Qi, YaoYang Liu et al.ICLR 2026 · 8 citations
- GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?Yiyang Zhou, Linjie Li, Shi Qiu, Zhengyuan Yang et al.EMNLP 2025 · 2 citations
- SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video UnderstandingZhenyu Yang, Yuhang Hu, Zemin Du, Dizhan Xue et al.ICLR 2025
- MVBench: A Comprehensive Multi-modal Video Understanding BenchmarkKunchang Li, Yali Wang, Yinan He, Yizhuo Li et al.CVPR 2024
