Enhancing Two Steps Textual Anomaly Detection through Anisotropy Mitigation
Pierre Fihey, Matthieu Labeau, Pavlo Mozharovskyi
摘要
Anomaly detection aims at distinguishing between in-distribution samples, which belong to the same distribution as the training set, and out-of-distribution samples, which lie outside of it. In textual anomaly detection, recent approaches routinely apply anomaly detection algorithms directly to embeddings extracted from pre-trained embedding models (two-stage approaches). However, the geometric properties of pre-trained embeddings can hinder the effectiveness of detection algorithms, which often rely on distance-based measures. In this work, we first highlight the relevance of similaritytrained models for textual anomaly detection. Beyond being trained to capture semantic similarities, these models also exhibit geometric properties that appear better suited to detection algorithms. We further demonstrate that, besides model choice, a simple post-processing step can significantly improve anomaly detection by adapting embeddings to the assumptions made by classical detection algorithms. The bulk of our experiments is done on a reformulation of the classification tasks from the MTEB benchmark into anomaly detection tasks 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary 等ACL 2020 · 被引用 539 次
- On the Sentence Embeddings from Pre-trained Language ModelsBohan Li, Hao Zhou, Junxian He, Mingxuan Wang 等EMNLP 2020 · 被引用 538 次
- LUNAR: Unifying Local Outlier Detection Methods via Graph Neural NetworksAdam Goodge, Bryan Hooi, See-Kiong Ng, Wee Siong NgAAAI 2022 · 被引用 144 次
相关 Paper
- Unsupervised Layer-Wise Score Aggregation for Textual OOD DetectionMaxime Darrin, Guillaume Staerman, Eduardo Dadalto Câmara Gomes, Jackie C. K. Cheung 等AAAI 2024 · 被引用 18 次
- Harnessing Large Language Models for Training-Free Video Anomaly DetectionLuca Zanella, Willi Menapace, Massimiliano Mancini, Yiming Wang 等CVPR 2024 · 被引用 57 次
- DE-CLIP: Few-Shot Anomaly Detection via Difference-Guided Embedding EditingYage Zhang, Yukun Jiang, Michael Backes, Yang ZhangACL 2026
- Back to Point: Exploring Point-Language Models for Zero-Shot 3D Anomaly DetectionKaiqiang Li, Gang Li, Mingle Zhou, Min Li 等CVPR 2026 · 被引用 2 次
- Hyperbolic Anomaly DetectionHuimin Li, Zhentao Chen, Yunhao Xu, Junlin HuCVPR 2024
