Enhancing Two Steps Textual Anomaly Detection through Anisotropy Mitigation
Pierre Fihey, Matthieu Labeau, Pavlo Mozharovskyi
Abstract
Anomaly detection aims at distinguishing between in-distribution samples, which belong to the same distribution as the training set, and out-of-distribution samples, which lie outside of it. In textual anomaly detection, recent approaches routinely apply anomaly detection algorithms directly to embeddings extracted from pre-trained embedding models (two-stage approaches). However, the geometric properties of pre-trained embeddings can hinder the effectiveness of detection algorithms, which often rely on distance-based measures. In this work, we first highlight the relevance of similaritytrained models for textual anomaly detection. Beyond being trained to capture semantic similarities, these models also exhibit geometric properties that appear better suited to detection algorithms. We further demonstrate that, besides model choice, a simple post-processing step can significantly improve anomaly detection by adapting embeddings to the assumptions made by classical detection algorithms. The bulk of our experiments is done on a reformulation of the classification tasks from the MTEB benchmark into anomaly detection tasks 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on8
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang et al.EMNLP 2020 · 538 citations
- LUNAR: Unifying Local Outlier Detection Methods via Graph Neural NetworksAdam Goodge, Bryan Hooi, See-Kiong Ng, Wee Siong NgAAAI 2022 · 144 citations
Related papers
- Unsupervised Layer-Wise Score Aggregation for Textual OOD DetectionMaxime Darrin, Guillaume Staerman, Eduardo Dadalto Câmara Gomes, Jackie C. K. Cheung et al.AAAI 2024 · 18 citations
- Harnessing Large Language Models for Training-Free Video Anomaly DetectionLuca Zanella, Willi Menapace, Massimiliano Mancini, Yiming Wang et al.CVPR 2024 · 57 citations
- DE-CLIP: Few-Shot Anomaly Detection via Difference-Guided Embedding EditingYage Zhang, Yukun Jiang, Michael Backes, Yang ZhangACL 2026
- Back to Point: Exploring Point-Language Models for Zero-Shot 3D Anomaly DetectionKaiqiang Li, Gang Li, Mingle Zhou, Min Li et al.CVPR 2026 · 2 citations
- Hyperbolic Anomaly DetectionHuimin Li, Zhentao Chen, Yunhao Xu, Junlin HuCVPR 2024
