SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM
Ming Nie, Dan Ding, Chunwei Wang, Yuanfan Guo, Jianhua Han, Hang Xu, Li Zhang
Abstract
Large language models (LLMs) have demonstrated exceptional capabilities in text understanding, which has paved the way for their expansion into video LLMs (Vid-LLMs) to analyze video data. However, current Vid-LLMs struggle to simultaneously retain high-quality frame-level semantic information (i.e., a sufficient number of tokens per frame) and comprehensive video-level temporal information (i.e., an adequate number of sampled frames per video). This limitation hinders the advancement of Vid-LLMs towards fine-grained video understanding. To address this issue, we introduce the SlowFocus mechanism, which significantly enhances the equivalent sampling frequency without compromising the quality of frame-level visual tokens. SlowFocus begins by identifying the query-related temporal segment based on the posed question, then performs dense sampling on this segment to extract local high-frequency features. A multi-frequency mixing attention module is further leveraged to aggregate these local high-frequency details with global low-frequency contexts for enhanced temporal comprehension. Additionally, to tailor Vid-LLMs to this innovative mechanism, we introduce a set of training strategies aimed at bolstering both temporal grounding and detailed temporal reasoning capabilities. Furthermore, we establish FineAction-CGR, a benchmark specifically devised to assess the ability of Vid-LLMs to process fine-grained temporal understanding tasks. Comprehensive experiments demonstrate the superiority of our mechanism across both existing public video understanding benchmarks and our proposed FineAction-CGR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 15d36a50-1dbf-4697-a2f3-5aa6e349e5e4Cited by top-tier papers8
- Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video ReasoningHaoji Zhang, Xin Gu, Jiawen Li, Chixiang Ma et al.CVPR 2026 · 92 citations
- FluxMem: Adaptive Hierarchical Memory for Streaming Video UnderstandingYiweng Xie, Bo He, Junke Wang, Xiangyu Zheng et al.CVPR 2026 · 25 citations
- Vid-SME: Membership Inference Attacks against Large Video Understanding ModelsQi Li, Runpeng Yu, Xinchao WangNeurIPS 2025 · 19 citations
- Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision EncodersAli Rasekh, Erfan Bagheri Soula, Omid Daliran, Simon Gottschalk et al.NeurIPS 2025 · 10 citations
- Improve Temporal Reasoning in Multimodal Large Language Models via Video Contrastive DecodingDaiqing Qi, Dongliang Guo, Hanzhang Yuan, Handong Zhao et al.NeurIPS 2025 · 5 citations
Builds on12
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li et al.ICLR 2024 · 467 citations
Related papers
- GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal GroundingRong Fan, Kaiyan Xiao, Minghao Zhu, Liuyi Wang et al.CVPR 2026 · 1 citation
- VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video ReasoningYang Ding, Xin Lai, Yizhen Zhang, Wei Li et al.ICLR 2026 · 26 citations
- LongVU: Spatiotemporal Adaptive Compression for Long Video-Language UnderstandingXiaoqian Shen, Yunyang Xiong, Changsheng Zhao, Lemeng Wu et al.ICML 2025
- Divid: Disentangled Spatial-Temporal Modeling within LLMs for Temporally Grounded Video UnderstandingYepeng Tang, Weining Wang, Longteng Guo, Tongtian Yue et al.ICLR 2026
- SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding CapabilityJiankang Wang, Zhihan Zhang, Zhihang Liu, Yang Li et al.AAAI 2026 · 20 citations
