Lune

CVPR2026顶会

Alert-CLIP: Abnormality-aware Latent-Enhanced Representation Tuning of CLIP for Video Anomaly Detection

Yiyan Zhu, Menghao Zhang, Haifeng Sun, Pengfei Ren, Xianao Chu, Chenye Xu, Hong Tan, Jinghan Wang, Qi Qi, Jingyu Wang

出版方
2026年份

摘要

With the rise of pre-trained vision-language models such as CLIP, performing video anomaly detection (VAD) through cross-modal reasoning has become an emerging trend. However, we observe that CLIP still suffers from weak abnormality awareness: normal and abnormal descriptions are highly entangled in the text embedding space, causing video features to assign nearly indistinguishable similarity scores to both types of prompts. To address this issue, we propose Alert-CLIP, an abnormality-aware latentenhanced tuning framework that tailors CLIP for VAD. Alert-CLIP introduces a multi-level alignment strategy:

(1) video-label alignment, which reshapes the semantic space to establish a coarse-level foundation for abnormality awareness; (2) region-text alignment, which explicitly associates anomaly-related regions with detailed descriptions to strengthen fine-grained perception; (3) region-semantic alignment, which contrasts anomalous regions against multiple hard negative samples to enhance abnormality-aware discrimination. To support this training, we construct VAGTA. Extensive experiments show that Alert-CLIP consistently surpasses CLIP across weakly supervised, zeroshot, and open-vocabulary settings. VAGTA is publicly

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper25

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖