Memory Matters: Boosting Training-Free Zero-Shot Temporal Action Localization with a Learnable Lookup Table
Han Jiang, Haoyu Tang, Xiaoxuan Mu, Chen Li, Jihua Zhu
Abstract
Zero-Shot Temporal Action Localization (ZS-TAL) aims to classify and localize actions in untrimmed videos that are unseen during training. Existing training-based ZS-TAL methods typically rely on fine-tuning models on large-scale annotated training data. This can be impractical in realworld applications and damage its generalization. As a result, Training-Free ZS-TAL has gained attention, which directly leverages Vision-Language Models (VLM) to enable action localization without any additional training. However, current techniques perform test-time adaptation independently on each video, neglecting the potential benefit of accumulating knowledge from historical test videos.
To address this, we propose a learnable lookup table (LLT) framework. During testing, we continuously update the lookup table by incorporating high-confidence, diverse lookup candidates to construct an action-positive lookup item. Additionally, we introduce a learnable residual module to adapt the corresponding lookup item to the current video context features. Finally, we employ refined activation scores to select accurate video frames and further adjust the text prototypes. This simple yet effective textvisual collaboration enables training-free ZS-TAL to harness knowledge from historical videos. Extensive experiments show our method significantly outperforms state-ofthe-art zero-shot VLM baselines, validating the effectiveness of our framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc5aa213-2602-4b1b-91b2-efa8d9424ef4Builds on28
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen et al.ICLR 2021 · 1,731 citations
- Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language ModelsManli Shu, Weili Nie, De-An Huang, Zhiding Yu et al.NeurIPS 2022 · 603 citations
- VadCLIP: Adapting Vision-Language Models for Weakly Supervised Video Anomaly DetectionPeng Wu, Xuerong Zhou, Guansong Pang, Lingru Zhou et al.AAAI 2024 · 220 citations
- CLIPood: Generalizing CLIP to Out-of-DistributionsYang Shu, Xingzhuo Guo, Jialong Wu, Ximei Wang et al.ICML 2023 · 122 citations
- UnLoc: A Unified Framework for Video Localization TasksShen Yan, Xuehan Xiong, Arsha Nagrani, Anurag Arnab et al.ICCV 2023 · 82 citations
Related papers
- Test-Time Zero-Shot Temporal Action LocalizationBenedetta Liberatori, Alessandro Conti, Paolo Rota, Yiming Wang et al.CVPR 2024
- Distilling Vision-Language Pre-Training to Collaborate with Weakly-Supervised Temporal Action LocalizationChen Ju, Kunhao Zheng, Jinxiang Liu, Peisen Zhao et al.CVPR 2023
- Dynamic Multimodal Prototype Learning in Vision-Language ModelsXingyu Zhu, Shuo Wang, Beier Zhu, Miaoge Li et al.ICCV 2025
- MAtch, eXpand and Improve: Unsupervised Finetuning for Zero-Shot Action Recognition with Language KnowledgeWei Lin, Leonid Karlinsky, Nina Shvetsova, Horst Possegger et al.ICCV 2023 · 52 citations
- Towards Mitigating Modality Bias in Vision-Language Models for Temporal Action LocalizationJiaqi Li, Guangming Wang, Shuntian Zheng, Minzhe Ni et al.ACL 2026 · 1 citation
