SpikingVTG: A Spiking Detection Transformer for Video Temporal Grounding
Malyaban Bal, Brian Matejek, Susmit Jha, Adam D. Cobb
Abstract
Video Temporal Grounding (VTG) aims to retrieve precise temporal segments in a video conditioned on natural language queries. Unlike conventional neural frameworks that rely heavily on computationally expensive dense matrix multi-plications, Spiking Neural Networks (SNNs)—previously underexplored in this domain—offer a unique opportunity to tackle VTG tasks through bio-plausible spike-based communication and an event-driven accumulation-based computational paradigm. We introduce SpikingVTG, a multi-modal spiking detection transformer, designed to harness the computational simplicity and sparsity of SNNs for VTG tasks. Leveraging the temporal dynamics of SNNs, our model introduces a Saliency Feedback Gating (SFG) mechanism that assigns dynamic saliency scores to video clips and applies multiplicative gating to highlight relevant clips while suppressing less informative ones. SFG enhances performance and reduces computational overhead by minimizing neural activity. We analyze the layer-wise convergence dynamics of SFG-enabled model and apply implicit differentiation at equilibrium to enable efficient, BPTT-free training. To improve generalization and maximize performance, we enable knowledge transfer by optimizing a Cos-L2 representation matching loss that aligns the layer-wise representation and attention maps of a non-spiking teacher with those of our student SpikingVTG. Additionally, we present Normalization-Free (NF)-SpikingVTG, which eliminates non-local operations like softmax and layer normalization, and an extremely quantized 1-bit (NF)-SpikingVTG variant for potential deployment on edge devices. Our models achieve competitive results on QVHighlights, Charades-STA, TACoS, and YouTube Highlights, establishing a strong baseline for multi-modal spiking VTG solutions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 33156782-94e2-4807-986a-e7e0cb27a9a5Builds on13
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural LanguageSongyang Zhang, Houwen Peng, Jianlong Fu, Jiebo LuoAAAI 2020 · 579 citations
- Detecting Moments and Highlights in Videos via Natural Language QueriesJie Lei, Tamara L. Berg, Mohit BansalNeurIPS 2021 · 425 citations
- UniVTG: Towards Unified Video-Language Temporal GroundingKevin Qinghong Lin, Pengchuan Zhang, Joya Chen, Shraman Pramanick et al.ICCV 2023 · 221 citations
- UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight DetectionYe Liu, Siyuan Li, Yang Wu, Chang Wen Chen et al.CVPR 2022 · 150 citations
Related papers
- Empower Words: DualGround for Structured Phrase and Sentence-Level Temporal GroundingMinseok Kang, Minhyeok Lee, Minjung Kim, Donghyeong Kim et al.NeurIPS 2025 · 4 citations
- Let Me Finish My Sentence: Video Temporal Grounding with Holistic Text UnderstandingJongbhin Woo, Hyeonggon Ryu, Youngjoon Jang, Jae-Won Cho et al.ACM MM 2024 · 3 citations
- Sparse-Dense Side-Tuner for Efficient Video Temporal GroundingDavid Pujol-Perich, Sergio Escalera, Albert ClapésICCV 2025 · 2 citations
- Spiking Transformers for Event-based Single Object TrackingJiqing Zhang, Bo Dong, Haiwei Zhang, Jianchuan Ding et al.CVPR 2022 · 171 citations
- STVGBert: A Visual-linguistic Transformer based Framework for Spatio-temporal Video GroundingRui Su, Qian Yu, Dong XuICCV 2021 · 75 citations
