Lune

CVPR2026顶会

Memory Matters: Boosting Training-Free Zero-Shot Temporal Action Localization with a Learnable Lookup Table

Han Jiang, Haoyu Tang, Xiaoxuan Mu, Chen Li, Jihua Zhu

出版方
2026年份

摘要

Zero-Shot Temporal Action Localization (ZS-TAL) aims to classify and localize actions in untrimmed videos that are unseen during training. Existing training-based ZS-TAL methods typically rely on fine-tuning models on large-scale annotated training data. This can be impractical in realworld applications and damage its generalization. As a result, Training-Free ZS-TAL has gained attention, which directly leverages Vision-Language Models (VLM) to enable action localization without any additional training. However, current techniques perform test-time adaptation independently on each video, neglecting the potential benefit of accumulating knowledge from historical test videos.

To address this, we propose a learnable lookup table (LLT) framework. During testing, we continuously update the lookup table by incorporating high-confidence, diverse lookup candidates to construct an action-positive lookup item. Additionally, we introduce a learnable residual module to adapt the corresponding lookup item to the current video context features. Finally, we employ refined activation scores to select accurate video frames and further adjust the text prototypes. This simple yet effective textvisual collaboration enables training-free ZS-TAL to harness knowledge from historical videos. Extensive experiments show our method significantly outperforms state-ofthe-art zero-shot VLM baselines, validating the effectiveness of our framework.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper28

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖