ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting
Yeonkyung Lee, Dayun Ju, Youngmin Kim, Seil Kang, Seong Jae Hwang
Abstract
Recent advancements in Video Large Language Models (VideoLLMs) have enabled strong performance across diverse multimodal video tasks. To reduce the high computational cost of processing dense video frames, efficiency-oriented methods such as frame selection have been widely adopted. While effective at minimizing redundancy, these methods often cause notable performance drops on tasks requiring temporal reasoning. Unlike humans, who can infer event progression from sparse visual cues, VideoLLMs frequently misinterpret temporal relations when intermediate frames are omitted. To address this limitation, we explore visual prompting (VP) as a lightweight yet effective way to enhance temporal understanding in VideoLLMs. Our analysis reveals that simply annotating each frame with explicit ordinal information helps the model perceive temporal continuity. This visual cue also supports frame-level referencing and mitigates positional ambiguity within a sparsely sampled sequence. Building on these insights, we introduce ViKey, a training-free framework that combines VP with a lightweight Keyword-Frame Mapping (KFM) module. KFM leverages frame indices as dictionary-like keys to link textual cues to the most relevant frames, providing explicit temporal anchors during inference. Despite its simplicity, our approach substantially improves temporal reasoning and, on some datasets, preserves dense-frame baseline performance with as few as 20% of frames.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5cdb266b-a6ba-4ca5-9a4d-a7bdcad8f0c1Builds on34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- What does CLIP know about a red circle? Visual prompt engineering for VLMsAleksandar Shtedritski, Christian Rupprecht, Andrea VedaldiICCV 2023 · 262 citations
- Fine-Grained Visual PromptingLingfeng Yang, Yueze Wang, Xiang Li, Xinlong Wang et al.NeurIPS 2023 · 129 citations
- LLaVA-Prumerge: Adaptive Token Reduction for Efficient Large Multimodal ModelsYuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee et al.ICCV 2025 · 37 citations
- Vista-llama: Reducing Hallucination in Video Language Models via Equal Distance to Visual TokensFan Ma, Xiaojie Jin, Heng Wang, Yuchen Xian et al.CVPR 2024 · 11 citations
Related papers
- Efficient Frame Selection for Long Video Understanding via Reinforcement LearningYaxuan Qin, Hefei Li, Wenqi Mu, Yancheng HeCVPR 2026 · 6 citations
- M-LLM Based Video Frame Selection for Efficient Video UnderstandingKai Hu, Feng Gao, Xiaohan Nie, Peng Zhou et al.CVPR 2025
- VideoITG: Multimodal Video Understanding with Instructed Temporal GroundingShihao Wang, Guo Chen, De-An Huang, Zhiqi Li et al.CVPR 2026 · 35 citations
- Improve Temporal Reasoning in Multimodal Large Language Models via Video Contrastive DecodingDaiqing Qi, Dongliang Guo, Hanzhang Yuan, Handong Zhao et al.NeurIPS 2025 · 5 citations
- VTimeCoT: Thinking by Drawing for Video Temporal Grounding and ReasoningJinglei Zhang, Yuanfan Guo, Rolandos Alexandros Potamias, Jiankang Deng et al.ICCV 2025 · 4 citations
