Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding
Zelin Zheng, Xinyan Liu, Ruixin Li, Antoni B. Chan, Guorong Li, Qingming Huang, Laiyun Qing
摘要
Current Video-LLM approaches for Video Temporal Grounding (VTG) typically rely on direct timestamp generation from an unstructured visual-token stream, often leading to brittle numerics and inconsistent boundaries. To address this, we propose Foresee-to-Ground (F2G), a framework that reformulates VTG as a verifiable Identify-then-Measure problem. F2G integrates Predictive Temporal Perception with Evidence-Driven Reasoning: it learns boundary-sensitive temporal representations to build a video-wide evidence pool of candidate event segments, and exposes these segments to the LLM as citable evidence units that bind boundary prediction to explicit event hypotheses. By decoupling event identification from precise boundary measurement, F2G stabilizes grounding and makes predictions verifiable. Extensive experiments demonstrate that F2G consistently improves grounding accuracy across diverse benchmarks, transfers robustly across different Video-LLM backbones, and preserves general video understanding capabilities. Our project is available at https://github.com/zelion2003/Foresee-to-Ground.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper25
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
- Detecting Moments and Highlights in Videos via Natural Language QueriesJie Lei, Tamara L. Berg, Mohit BansalNeurIPS 2021 · 被引用 425 次
- Fast Video Moment RetrievalJunyu Gao, Changsheng XuICCV 2021 · 被引用 132 次
- An Empirical Study of CLIP for Text-Based Person SearchMin Cao, Yang Bai, Ziyin Zeng, Mang Ye 等AAAI 2024 · 被引用 111 次
相关 Paper
- GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal GroundingRong Fan, Kaiyan Xiao, Minghao Zhu, Liuyi Wang 等CVPR 2026 · 被引用 1 次
- TRACE: Temporal Grounding Video LLM via Causal Event ModelingYongxin Guo, Jingyu Liu, Mingda Li, Qingbin Liu 等ICLR 2025
- Factorized Learning for Temporally Grounded Video-Language ModelsWenzheng Zeng, Difei Gao, Mike Zheng Shou, Hwee Tou NgICCV 2025 · 被引用 2 次
- VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal GroundingYongxin Guo, Jingyu Liu, Mingda Li, Dingxin Cheng 等AAAI 2025 · 被引用 27 次
- T2SGrid: Temporal-to-Spatial Gridification for Video Temporal GroundingChaohong Guo, Yihan He, Yongwei Nie, Fei Ma 等CVPR 2026 · 被引用 2 次
