Temporal Reasoning Transfer from Text to Video
Lei Li, Yuanxin Liu, Linli Yao, Peiyuan Zhang, Chenxin An, Lean Wang, Xu Sun, Lingpeng Kong, Qi Liu
摘要
Video Large Language Models (Video LLMs) have shown promising capabilities in video comprehension, yet they struggle with tracking temporal changes and reasoning about temporal relationships. While previous research attributed this limitation to the ineffective temporal encoding of visual inputs, our diagnostic study reveals that video representations contain sufficient information for even small probing classifiers to achieve perfect accuracy. Surprisingly, we find that the key bottleneck in Video LLMs' temporal reasoning capability stems from the underlying LLM's inherent difficulty with temporal concepts, as evidenced by poor performance on textual temporal question-answering tasks. Building on this discovery, we introduce the Textual Temporal reasoning Transfer (T3). T3 synthesizes diverse temporal reasoning tasks in pure text format from existing image-text datasets, addressing the scarcity of video samples with complex temporal scenarios. Remarkably, without using any video data, T3 enhances LongVA-7B's temporal understanding, yielding a 5.3 absolute accuracy improvement on the challenging TempCompass benchmark, which enables our model to outperform ShareGPT4Video-8B trained on 28,000 video samples. Additionally, the enhanced LongVA-7B model achieves competitive performance on comprehensive video benchmarks. For example, it achieves a 49.7 accuracy on the Temporal Reasoning task of Video-MME, surpassing powerful large-scale models such as InternVL-Chat-V1.5-20B and VILA1.5-40B. Further analysis reveals a strong correlation between textual and video temporal task performance, validating the efficacy of transferring temporal reasoning abilities from text to video domains.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement LearningLong Xing, Xiaoyi Dong, Yuhang Zang, Yuhang Cao 等ICLR 2026 · 被引用 37 次
- AVoCaDO: An Audiovisual Video Captioner Driven by Temporal OrchestrationXinlong Chen, Yue Ding, Weihong Lin, Jingyun Hua 等ICLR 2026 · 被引用 27 次
- VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?Yuanxin Liu, Kun Ouyang, Haoning Wu, Yi Liu 等ICLR 2026 · 被引用 20 次
- TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming VideosLinli Yao, Yicheng Li, Yuancheng Wei, Lei Li 等ACM MM 2025 · 被引用 14 次
- Mixture of Cognitive Reasoners: Modular Reasoning with Brain-Like SpecializationBadr AlKhamissi, C. Nicolò De Sabbata, Greta Tuckute, Zeming Chen 等ICLR 2026 · 被引用 12 次
相关 Paper
- Improve Temporal Reasoning in Multimodal Large Language Models via Video Contrastive DecodingDaiqing Qi, Dongliang Guo, Hanzhang Yuan, Handong Zhao 等NeurIPS 2025 · 被引用 5 次
- Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision EncodersAli Rasekh, Erfan Bagheri Soula, Omid Daliran, Simon Gottschalk 等NeurIPS 2025 · 被引用 10 次
- Map the Flow: Revealing Hidden Pathways of Information in VideoLLMsMinji Kim, Taekyung Kim, Bohyung HanICLR 2026 · 被引用 8 次
- MVBench: A Comprehensive Multi-modal Video Understanding BenchmarkKunchang Li, Yali Wang, Yinan He, Yizhuo Li 等CVPR 2024
- Momentor: Advancing Video Large Language Model with Fine-Grained Temporal ReasoningLong Qian, Juncheng Li, Yu Wu, Yaobo Ye 等ICML 2024 · 被引用 121 次
