DisTime: Distribution-Based Time Representation for Video Large Language Models
Yingsen Zeng, Zepeng Huang, Yujie Zhong, Chengjian Feng, Jie Hu, Lin Ma, Yang Liu
Abstract
Despite advances in general video understanding, Video Large Language Models (Video-LLMs) face challenges in precise temporal localization due to discrete time representations and limited temporally aware datasets. Existing methods for temporal expression either conflate time with text-based numerical values, add a series of dedicated temporal tokens, or regress time using specialized temporal grounding heads. To address these issues, we introduce DisTime, a lightweight framework designed to enhance temporal comprehension in Video-LLMs. DisTime employs a learnable token to create a continuous temporal embedding space and incorporates a Distribution-based Time Decoder that generates temporal probability distributions, effectively mitigating boundary ambiguities and maintaining temporal continuity. Additionally, the Distribution-based Time Encoder re-encodes timestamps to provide time markers for Video-LLMs. To overcome temporal granularity limitations in existing datasets, we propose an automated annotation paradigm that combines the captioning capabilities of Video-LLMs with the localization expertise of dedicated temporal models. This leads to the creation of InternVid-TG, a substantial dataset with 1.25M temporally grounded events across 179k videos, surpassing ActivityNet-Caption by 55 times. Extensive experiments demonstrate that DisTime achieves state-of-the-art performance across benchmarks in three time-sensitive tasks while maintaining competitive performance in Video QA tasks. Code and data are released at https://github.com/josephzpng/DisTime.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 594aa7e6-cb8a-4878-9553-911990989a16Cited by top-tier papers6
- TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMsJun Zhang, Teng Wang, Yuying Ge, Yixiao Ge et al.CVPR 2026 · 48 citations
- RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous DrivingZhijian Huang, Chengjian Feng, Feng Yan, Baihui Xiao et al.ICCV 2025 · 6 citations
- OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal GroundingMinghang Zheng, Zihao Yin, Yi Yang, Yuxin Peng et al.CVPR 2026 · 4 citations
- UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language ModelsHewen Pan, Cong Wei, Dashuang Liang, Zepeng Huang et al.CVPR 2026 · 3 citations
- Temporal-Aware Reasoning Optimization for Video Temporal GroundingMinghang Zheng, Zihao Yin, YI YANG, Yuxin Peng et al.ICML 2026
Builds on22
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object DetectionXiang Li, Wenhai Wang, Lijun Wu, Shuo Chen et al.NeurIPS 2020 · 2,118 citations
Related papers
- VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal GroundingYongxin Guo, Jingyu Liu, Mingda Li, Dingxin Cheng et al.AAAI 2025 · 27 citations
- GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal GroundingRong Fan, Kaiyan Xiao, Minghao Zhu, Liuyi Wang et al.CVPR 2026 · 1 citation
- Universal Video Temporal Grounding with Generative Multi-modal Large Language ModelsZeqian Li, Shangzhe Di, Zhonghua Zhai, Weilin Huang et al.NeurIPS 2025 · 30 citations
- Divid: Disentangled Spatial-Temporal Modeling within LLMs for Temporally Grounded Video UnderstandingYepeng Tang, Weining Wang, Longteng Guo, Tongtian Yue et al.ICLR 2026
- VTimeLLM: Empower LLM to Grasp Video MomentsBin Huang, Xin Wang, Hong Chen, Zihan Song et al.CVPR 2024
