LLM as an Algorithmist: Enhancing Anomaly Detectors via Programmatic Synthesis
Hangting Ye, Jinmeng Li, He Zhao, Mingchen Zhuge, Dandan Guo, Yi Chang, Hongyuan Zha
摘要
Existing anomaly detection (AD) methods for tabular data usually rely on some assumptions about anomaly patterns, leading to inconsistent performance in realworld scenarios. While Large Language Models (LLMs) show remarkable reasoning capabilities, their direct application to tabular AD is impeded by fundamental challenges, including difficulties in processing heterogeneous data and significant privacy risks. To address these limitations, we propose LLM-DAS, a novel framework that repositions the LLM from a "data processor" to an "algorithmist". Instead of being exposed to raw data, our framework leverages the LLM's ability to reason about algorithms. It analyzes a high-level description of a given detector to understand its intrinsic weaknesses and then generates detector-specific, dataagnostic Python code to synthesize "hard-to-detect" anomalies that exploit these vulnerabilities. This generated synthesis program, which is reusable across diverse datasets, is then instantiated to augment training data, systematically enhancing the detector's robustness by transforming the problem into a more discriminative two-class classification task. Extensive experiments on 36 TAD benchmarks show that LLM-DAS consistently boosts the performance of mainstream detectors. By bridging LLM reasoning with classic AD algorithms via programmatic synthesis, LLM-DAS offers a scalable, effective, and privacy-preserving approach to patching the logical blind spots of existing detectors. The source code is available at https://github.com/HangtingYe/LLM_DAS# .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 被引用 1,847 次
- Classification-Based Anomaly Detection for General DataLiron Bergman, Yedid HoshenICLR 2020 · 被引用 412 次
- Learning and Evaluating Representations for Deep One-Class ClassificationKihyuk Sohn, Chun-Liang Li, Jinsung Yoon, Minho Jin 等ICLR 2021 · 被引用 243 次
- LIFT: Language-Interfaced Fine-Tuning for Non-language Machine Learning TasksTuan Dinh, Yuchen Zeng, Ruisu Zhang, Ziqian Lin 等NeurIPS 2022 · 被引用 222 次
相关 Paper
- AnoLLM: Large Language Models for Tabular Anomaly DetectionChe-Ping Tsai, Ganyu Teng, Phillip Wallis, Wei DingICLR 2025
- MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language ModelsXincheng Yao, Zefeng Qian, Chao Shi, Jiayang Song 等CVPR 2026 · 被引用 2 次
- ReTabAD: A Benchmark for Restoring Semantic Context in Tabular Anomaly DetectionSanghyu Yoon, Dongmin Kim, Suhee Yoon, Ye Seul Sim 等ICLR 2026 · 被引用 3 次
- RealVul: Can We Detect Vulnerabilities in Web Applications with LLM?Di Cao, Yong Liao, Xiuwei ShangEMNLP 2024 · 被引用 11 次
- Label Annotation for Tabular Anomaly Detection with Large Language ModelsHaihong Zhao, Aochuan Chen, Miao Peng, Xiaolong Fan 等KDD 2026
