HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
Joungbin An, Kristen Grauman
摘要
Video temporal grounding, the task of localizing the start and end times of a natural language query in untrimmed video, requires capturing both global context and finegrained temporal detail. This challenge is particularly pronounced in long videos, where existing methods often compromise temporal fidelity by over-downsampling or using fixed windows. We present HieraMamba, a hierarchical architecture that preserves temporal structure and semantic richness across scales. At its core are Anchor-MambaPooling (AMP) blocks, which utilize Mamba's selective scanning to produce compact anchor tokens summarizing video content across scales. We further introduce anchor-conditioned and segment-pooled contrastive losses-two complementary objectives that encourage anchors to retain local detail while remaining globally discriminative. HieraMamba sets a new state-of-the-art on Ego4D-NLQ, MAD, and TACoS, demonstrating precise, temporally faithful localization in long, untrimmed videos. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper49
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Distance-IoU Loss: Faster and Better Learning for Bounding Box RegressionZhaohui Zheng, Ping Wang, Wei Liu, Jinze Li 等AAAI 2020 · 被引用 4,823 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
相关 Paper
- HiTeA: Hierarchical Temporal Alignment for Training-Free Long-Video Temporal GroundingXinyi Xu, Hongsong Wang, Guo-Sen Xie, Caifeng Shan 等ICLR 2026
- Multi-Scale Contrastive Learning for Video Temporal GroundingThong Thanh Nguyen, Yi Bin, Xiaobao Wu, Zhiyuan Hu 等AAAI 2025 · 被引用 7 次
- CONE: An Efficient COarse-to-fiNE Alignment Framework for Long Video Temporal GroundingZhijian Hou, Wanjun Zhong, Lei Ji, Difei Gao 等ACL 2023 · 被引用 17 次
- Universal Video Temporal Grounding with Generative Multi-modal Large Language ModelsZeqian Li, Shangzhe Di, Zhonghua Zhai, Weilin Huang 等NeurIPS 2025 · 被引用 30 次
- Measure Twice, Cut Once: A Semantic-Oriented Approach to Video Temporal Localization with Video LLMsZongshang Pang, Mayu Otani, Yuta NakashimaICLR 2026 · 被引用 2 次
