Improving LLM Video Understanding with 16 Frames Per Second
Yixuan Li, Changli Tang, Jimin Zhuang, Yudong Yang, Guangzhi Sun, Wei Li, Zejun Ma, Chao Zhang
Abstract
Human vision is dynamic and continuous. However, in video understanding with multimodal large language models (LLMs), existing methods primarily rely on static features extracted from images sampled at a fixed low frame rate of frame-per-second (FPS) ⩽2, leading to critical visual information loss. In this paper, we introduce F-16, the first multimodal LLM designed for high-frame-rate video understanding. By increasing the frame rate to 16 FPS and compressing visual tokens within each 1-second clip, F-16 efficiently captures dynamic visual features while preserving key semantic information. Experimental results demonstrate that higher frame rates considerably enhance video understanding across multiple benchmarks, providing a new approach to improving video LLMs beyond scaling model size or training data. F-16 achieves state-ofthe-art performance among 7-billion-parameter video LLMs on both general and fine-grained video understanding benchmarks, such as Video-MME and TemporalBench. Furthermore, F-16 excels in complex spatiotemporal tasks, including high-speed sports analysis (e.g., basketball, football, gymnastics, and diving), outperforming SOTA proprietary visual models like GPT-4o and Gemini-1.5-pro. Additionally, we introduce a novel decoding method for F-16 that enables highly efficient low-frame-rate inference without requiring model retraining. We will release the source code, model checkpoints, and data at https://github.com/bytedance/F-16 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc05a2d6-2146-44a5-a5b0-ffd0445bb0e3Cited by top-tier papers6
- LongVT: Incentivizing "Thinking with Long Videos" via Native Tool CallingZuhao Yang, Sudong Wang, Kaichen Zhang, Keming Wu et al.CVPR 2026 · 63 citations
- VideoNSA: Native Sparse Attention Scales Video UnderstandingEnxin Song, Wenhao Chai, Shusheng Yang, Ethan Armand et al.ICLR 2026 · 11 citations
- WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLMChangli Tang, Qinfan Xiao, Ke Mei, Tianyi Wang et al.ICLR 2026 · 9 citations
- Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive GazingBaifeng Shi, Stephanie Fu, Long Lian, Hanrong Ye et al.CVPR 2026 · 9 citations
- video-SALMONN S: Memory-Enhanced Streaming Audio-Visual LLMGuangzhi Sun, Yixuan Li, Xiaodong Wu, Yudong Yang et al.ICML 2026 · 2 citations
Builds on12
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Streaming Long Video Understanding with Large Language ModelsRui Qian, Xiaoyi Dong, Pan Zhang, Yuhang Zang et al.NeurIPS 2024 · 216 citations
Related papers
- Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional TokenizationYang Jin, Zhicheng Sun, Kun Xu, Kun Xu et al.ICML 2024 · 94 citations
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video AnalysisChaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li et al.CVPR 2025
- Q-Frame: Query-Aware Frame Selection and Multi-Resolution Adaptation for Video-LLMsShaojie Zhang, Jiahui Yang, Jianqin Yin, Zhenbo Luo et al.ICCV 2025 · 15 citations
- B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal TokensZhuqiang Lu, Zhenfei Yin, Mengwei He, Zhihui Wang et al.ICCV 2025 · 3 citations
- DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video UnderstandingXiaoyi Bao, Chenwei Xie, Hao Tang, Tingyu Weng et al.ICCV 2025 · 6 citations
