ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting
Yeonkyung Lee, Dayun Ju, Youngmin Kim, Seil Kang, Seong Jae Hwang
摘要
Recent advancements in Video Large Language Models (VideoLLMs) have enabled strong performance across diverse multimodal video tasks. To reduce the high computational cost of processing dense video frames, efficiency-oriented methods such as frame selection have been widely adopted. While effective at minimizing redundancy, these methods often cause notable performance drops on tasks requiring temporal reasoning. Unlike humans, who can infer event progression from sparse visual cues, VideoLLMs frequently misinterpret temporal relations when intermediate frames are omitted. To address this limitation, we explore visual prompting (VP) as a lightweight yet effective way to enhance temporal understanding in VideoLLMs. Our analysis reveals that simply annotating each frame with explicit ordinal information helps the model perceive temporal continuity. This visual cue also supports frame-level referencing and mitigates positional ambiguity within a sparsely sampled sequence. Building on these insights, we introduce ViKey, a training-free framework that combines VP with a lightweight Keyword-Frame Mapping (KFM) module. KFM leverages frame indices as dictionary-like keys to link textual cues to the most relevant frames, providing explicit temporal anchors during inference. Despite its simplicity, our approach substantially improves temporal reasoning and, on some datasets, preserves dense-frame baseline performance with as few as 20% of frames.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- What does CLIP know about a red circle? Visual prompt engineering for VLMsAleksandar Shtedritski, Christian Rupprecht, Andrea VedaldiICCV 2023 · 被引用 262 次
- Fine-Grained Visual PromptingLingfeng Yang, Yueze Wang, Xiang Li, Xinlong Wang 等NeurIPS 2023 · 被引用 129 次
- LLaVA-Prumerge: Adaptive Token Reduction for Efficient Large Multimodal ModelsYuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee 等ICCV 2025 · 被引用 37 次
- Vista-llama: Reducing Hallucination in Video Language Models via Equal Distance to Visual TokensFan Ma, Xiaojie Jin, Heng Wang, Yuchen Xian 等CVPR 2024 · 被引用 11 次
相关 Paper
- Efficient Frame Selection for Long Video Understanding via Reinforcement LearningYaxuan Qin, Hefei Li, Wenqi Mu, Yancheng HeCVPR 2026 · 被引用 6 次
- M-LLM Based Video Frame Selection for Efficient Video UnderstandingKai Hu, Feng Gao, Xiaohan Nie, Peng Zhou 等CVPR 2025
- VideoITG: Multimodal Video Understanding with Instructed Temporal GroundingShihao Wang, Guo Chen, De-An Huang, Zhiqi Li 等CVPR 2026 · 被引用 35 次
- Improve Temporal Reasoning in Multimodal Large Language Models via Video Contrastive DecodingDaiqing Qi, Dongliang Guo, Hanzhang Yuan, Handong Zhao 等NeurIPS 2025 · 被引用 5 次
- VTimeCoT: Thinking by Drawing for Video Temporal Grounding and ReasoningJinglei Zhang, Yuanfan Guo, Rolandos Alexandros Potamias, Jiankang Deng 等ICCV 2025 · 被引用 4 次
