Lune

CVPR2026Top-tier venue

Alert-CLIP: Abnormality-aware Latent-Enhanced Representation Tuning of CLIP for Video Anomaly Detection

Yiyan Zhu, Menghao Zhang, Haifeng Sun, Pengfei Ren, Xianao Chu, Chenye Xu, Hong Tan, Jinghan Wang, Qi Qi, Jingyu Wang

2026Year

Abstract

With the rise of pre-trained vision-language models such as CLIP, performing video anomaly detection (VAD) through cross-modal reasoning has become an emerging trend. However, we observe that CLIP still suffers from weak abnormality awareness: normal and abnormal descriptions are highly entangled in the text embedding space, causing video features to assign nearly indistinguishable similarity scores to both types of prompts. To address this issue, we propose Alert-CLIP, an abnormality-aware latentenhanced tuning framework that tailors CLIP for VAD. Alert-CLIP introduces a multi-level alignment strategy:

(1) video-label alignment, which reshapes the semantic space to establish a coarse-level foundation for abnormality awareness; (2) region-text alignment, which explicitly associates anomaly-related regions with detailed descriptions to strengthen fine-grained perception; (3) region-semantic alignment, which contrasts anomalous regions against multiple hard negative samples to enhance abnormality-aware discrimination. To support this training, we construct VAGTA. Extensive experiments show that Alert-CLIP consistently surpasses CLIP across weakly supervised, zeroshot, and open-vocabulary settings. VAGTA is publicly

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 3367d773-1a42-48b6-a23b-04979185dff4

Builds on25

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines